internal Docker-free suite (local_lift), pytest-graded — NOT SWE-bench/Terminal-Bench
Significant
Measured on our own Docker-free suite, graded by pytest — one model, small self-contained Python tasks. This is not SWE-bench, and it does not generalise to real repositories.
- Baseline
- 48%
- With Chimera
- 71%
- Difference
- +23pp
- 95% confidence interval
- [+12.6%, +28.6%]
- Sample
- n = 100
- Model
- openrouter/mistralai/mistral-small-3.2-24b-instruct
A cheap weak model + Chimera's retry loop vs the cheap model alone, on a pre-registered n=100 suite (design + tasks committed before any model call). SIGNIFICANT: the paired 95% CI excludes zero. The lift is 28 tasks the loop RECOVERED (raw fail → verified pass) against 5 regressions. This is the RE-VERIFIED run: an earlier run of the same suite (9% → 15%) was graded with a test file the agent under test could edit — and on re-run it did edit one — so grading was hardened to restore the pristine test, and the lift replicated LARGER, not smaller. The superseded run is kept unedited beside the erratum in bench/local_lift/RESULTS.md. One model, one seed/task, small self-contained Python tasks — NOT SWE-bench, does not generalise to real repos. One run, no re-roll.