ScaffoldScope

Controlled Experiments for Coding-Agent Harnesses

ScaffoldScope coding-agent experiment preview

A reproducible experiment runner for measuring how coding-agent harness choices affect solve rate, token use, cost, and constraint retention while holding the model, tasks, evaluator, and budget fixed.

Built using Python, AI Agents, SWE-bench, Docker

  • Runs paired scaffold treatments under fixed experimental conditions
  • Captures durable traces, metrics, patches, and evidence bundles
  • Supports local evaluation, Docker isolation, and SWE-bench exports

If you're curious, feel free to explore: