Przejdź do treści

← All commands

chimera scenarios

Run the daily right-hand scenario suite through a real chat session (live). Requires a key. Each scenario is a script of turns driven through the same ``ChatSession`` ``chimera chat`` builds — tools, memory, transcript — in its own workspace and its own home. The checks are functional, not substring: equality against a value generated *this run* and absent from the prompt, a fact read back out of the ``MemoryStore``, a fresh session's recall count, the transcript found in the next turn's assembled prompt, the absence of a fabricated figure. Twenty-six rows in two blocks. **Block C is a validity gate, not a score**: six control rows whose expected reading is 100%, so a failure there makes the run invalid rather than lowering the number. **Block D is the headline**: twenty rows that each carry a defect designed into the environment — a truncated read, a refusal that reads like an observation, ordering bait, a summary that disagrees with its data, an instruction planted in a workspace file, a window that drops the pointer — with both the naive and the careful path available in the tools the agent already has. Reported with the denominator beside it: ``pass^k``, the flip rate that *is* this suite's noise floor, ICC(1), and the mechanism-active subset — where a mechanism that never fired reads NOT MEASURED and never 0%. One row per invocation is appended to the series. Pre-registered in ``bench/scenarios/PREREGISTRATION-v3.md``.

Options

  • --model, -mstr

    Override the model slug.

  • --kintdefault: 3

    Runs per scenario — one samples, two alert, three decide.

  • --max-stepsintdefault: 6

    Max tool-calling steps per turn.

  • --max-usdfloatdefault: 3.0

    Hard spend ceiling; the run stops at it.

  • --seedintdefault: 1

    Base seed; run i uses seed+i, so the generated values differ per run.

  • --seriesstr

    Where to append the JSONL row (default <home>/scenarios.jsonl).