
A reproducible experiment runner for measuring how coding-agent harness choices affect solve rate, token use, cost, and constraint retention while holding the model, tasks, evaluator, and budget fixed.
Built using Python, AI Agents, SWE-bench, Docker
If you're curious, feel free to explore: