internal Docker-free suite (local_lift), pytest-graded — NOT SWE-bench/Terminal-Bench
Значимо
Измерено на нашем собственном наборе без Docker, оценка через pytest — одна модель, небольшие самодостаточные задачи на Python. Это не SWE-bench, и на настоящие репозитории это не переносится.
- База
- 48%
- С Chimera
- 71%
- Разница
- +23pp
- 95% доверительный интервал
- [+12.6%, +28.6%]
- Выборка
- n = 100
- Модель
- openrouter/mistralai/mistral-small-3.2-24b-instruct
A cheap weak model + Chimera's retry loop vs the cheap model alone, on a pre-registered n=100 suite (design + tasks committed before any model call). SIGNIFICANT: the paired 95% CI excludes zero. The lift is 28 tasks the loop RECOVERED (raw fail → verified pass) against 5 regressions. This is the RE-VERIFIED run: an earlier run of the same suite (9% → 15%) was graded with a test file the agent under test could edit — and on re-run it did edit one — so grading was hardened to restore the pristine test, and the lift replicated LARGER, not smaller. The superseded run is kept unedited beside the erratum in bench/local_lift/RESULTS.md. One model, one seed/task, small self-contained Python tasks — NOT SWE-bench, does not generalise to real repos. One run, no re-roll.