Перейти к содержимому

Доказательства

Каждое число, которое публикует этот проект, включая те, где он проиграл.

Эти цифры прочитаны из снимков, которые продукт создаёт и штампует при каждом выпуске. Ничто на этой странице не набрано вручную — поэтому ничто здесь не может незаметно устареть.

Описана версия 0.43.0

internal Docker-free suite (local_lift), pytest-graded — NOT SWE-bench/Terminal-Bench

Значимо

Измерено на нашем собственном наборе без Docker, оценка через pytest — одна модель, небольшие самодостаточные задачи на Python. Это не SWE-bench, и на настоящие репозитории это не переносится.

База
48%
С Chimera
71%
Разница
+23pp
95% доверительный интервал
[+12.6%, +28.6%]
Выборка
n = 100
Модель
openrouter/mistralai/mistral-small-3.2-24b-instruct

A cheap weak model + Chimera's retry loop vs the cheap model alone, on a pre-registered n=100 suite (design + tasks committed before any model call). SIGNIFICANT: the paired 95% CI excludes zero. The lift is 28 tasks the loop RECOVERED (raw fail → verified pass) against 5 regressions. This is the RE-VERIFIED run: an earlier run of the same suite (9% → 15%) was graded with a test file the agent under test could edit — and on re-run it did edit one — so grading was hardened to restore the pristine test, and the lift replicated LARGER, not smaller. The superseded run is kept unedited beside the erratum in bench/local_lift/RESULTS.md. One model, one seed/task, small self-contained Python tasks — NOT SWE-bench, does not generalise to real repos. One run, no re-roll.

Источник: bench/local_lift/_reverify_n100/paired.json

SWE-bench Verified (41-instance django easy slice, out-of-sample replication)

Не значимо

Намеренно лёгкий срез SWE-bench Verified из одного репозитория. Это не оценка по SWE-bench Verified — для настоящей нужны все 500 случаев. Сама по себе разница не значима.

База
34.1%
С Chimera
43.9%
Разница
+9.8pp
95% доверительный интервал
[-3.5%, +16.7%]
Выборка
n = 41
Модель
openrouter/deepseek/deepseek-chat-v3.1

NOT a SWE-bench Verified score — a deliberately easy, single-repo slice; a real score needs the full 500. Graded only by the official swebench harness in Docker, never self-reported. This is the OUT-OF-SAMPLE replication: 41 instances never used by the earlier run, nothing else changed. The delta is NOT significant on its own (CI includes 0); pooling it with the earlier 19 gives +11.7% [+0.8%, +16.4%], which IS significant but mixes seen with unseen data and was pre-registered as secondary. DECOMPOSITION (4th run, middle arm restored on the same 41): scaffold alone +4.9% over baseline, the diff-gate a further +4.9% — BOTH components contribute in roughly equal halves, neither significant alone. That CONTRADICTED our own registered prediction that the scaffold would carry most of it, and withdrew an earlier reading that the gate was not what produced the gain; the retraction is in RESULTS.md. The tidy additivity is NOT claimed as a measured 50/50 split — each comparison rests on 5-6 discordant pairs. All three arms edit at the SAME rate (27-28 patches of 41); what climbs is precision, 50% -> 59% -> 67%. An earlier run of the same design at a starved step budget scored an exact 0.0pp and is published unchanged.

Источник: bench/swe_bench/RESULTS.md

Terminal-Bench (terminal-bench-core 0.1.1)

Не значимо

Здесь каркас не помог, и запуск опубликован как измерен. Оба плеча стоят у самого пола, и разница не значима.

База
7.5%
С Chimera
2.5%
Разница
-5pp
95% доверительный интервал
[-5%, +1.6%]
Выборка
n = 40
Модель
openrouter/deepseek/deepseek-chat-v3.1

Single-attempt, N=40, same model both arms. The scaffold did NOT lift here — deepseek-v3.1 is already competent, not the weak 'goldilocks' regime where scaffolding helps; both arms sit near the floor and the difference is not significant (CI includes 0). Published as measured; the point estimate is dominated by run-to-run variance at this floor.

Источник: bench/terminal_bench/RESULTS.md

Покрытие поверхностей

Покрытие внутренних поверхностей — какая часть кодовой базы подтверждена настоящими тестами. Это не утверждение о продукте: продукт находится в стадии альфа.

подтверждено критериев: 37 из 37

  • fusion5/5
  • evolution8/8
  • governance7/7
  • memory4/4
  • benchmarks5/5
  • resilience6/6
  • interop2/2

Отозванное утверждение

The agent extracts a skill from every verified success and reads them back on later runs. Whether that accumulation makes it measurably better at new tasks was tested across seven pre-registered runs. The sixth found a positive; the seventh, with more power, cut it to nothing.

Отозвано. Семь заранее объявленных запусков; единственный положительный результат не пережил воспроизведение с большей мощностью, поэтому утверждение сняли, а не оставили.