Przejdź do treści

Every command

Ta strona nie została jeszcze przetłumaczona, więc czytasz angielski oryginał.

Generated from the CLI itself, so it cannot describe a command that does not exist or miss one that does. Thirty-three of these appeared in no README and no doc before this page; a reference written by hand fixes that once and then goes stale in silence, which is the failure worth designing out.

Subcommands of a group are listed under their full path (agents list, cron add). Run chimera <command> --help for the full text of any entry.

Command What it does
a2a-card Print Chimera's A2A Agent Card JSON (serve it at /.well-known/agent.json).
acp Serve Chimera to an editor over the Agent Client Protocol (stdio).
agent Run the ReAct agent loop with native tools. Requires a provider key.
agents The agents you dispatch work to — as distinct from the one you converse with.
agents list Show the registry.
agents rm Forget an agent. Cards already filed under its lane are left exactly where they are.
agents set Add an agent, or replace the one with this id.
app Run the Chimera Desktop app: the HTTP+SSE API + the built React UI (needs the 'desktop' extra).
approve Answer a decision the agent is waiting on, from anywhere.
assist Your daily-driver assistant: cheap by default, escalates when it must.
bench Run the continuous-evolution benchmark on a demo task set. Requires a key.
bench-compare Report the honest A/B delta (+95% CI) between two benchmark result files.
brief Morning brief: parallel topic research through the hierarchy, one synthesized digest.
cascade-bench Four-arm bench: weak-only vs mid-only vs cascade vs fusion. Calls real models.
chat Interactive multi-turn chat — your terminal right-hand. Requires a key.
code Continue a desktop Code conversation from this terminal, through the running app.
code list The desktop app's Code conversations, newest first, with the id resume takes.
code resume Continue a desktop Code conversation here; the app runs the turn and shows it too.
context-curve Did runs carrying more context do worse? Measured on THIS machine's own logs.
crew Run a multi-agent crew on a task (Tier 3). Requires a provider key.
crew-isolated Tier-3: tool-using workers attempt ONE task, each in its own git worktree, verify-gated.
cron Manage scheduled jobs (crons and event SOPs).
cron add Add a cron, event- or webhook-triggered job.
cron disable Disable a job without deleting it.
cron doctor Ask the schedule what it is not telling you: what never ran, and what ran and lost.
cron enable Enable a job (e.g. an agent-proposed one) and schedule its next run.
cron fire Run every job registered for an event.
cron kill Stop a job's running (or next) dispatch — one run, not the schedule.
cron learn Propose crons from recurring tasks and create the ones you confirm.
cron list List scheduled jobs.
cron remove Remove a scheduled job by id.
decide Ask typed questions — yes/no, a choice, a score — and get probabilities back.
decisions Typed decisions: which model answers them, the log of what they answered, labels, a report and a refit.
decisions label Say what was true for one answer. A later label for the same id replaces an earlier one.
decisions log The latest answers, newest last, with their label when one was given.
decisions models The System One models OpenRouter lists, which one is active, and which carry a calibration map.
decisions refit Fit this deployment's own map on its labelled answers — pooled with the shipped rows while it
decisions report What the log holds: availability, the review budget, label coverage, and — where labels exist —
decisions use Choose the backend (and model) that answers typed decisions — written to .env in this folder,
delegations Measured vs counterfactual across delegations — what the hierarchy actually saved.
deliver Deliverable Mode: produce a polished, self-contained artifact. Requires a key.
doctor Check the environment and configuration. With --fix, repair safe setup issues.
drift Drift gate: check the workspace against a spec (Spec Growth). Exit 1 on drift.
evoclaw Stress-test continuous-evolution degradation: naive vs guarded. Requires a key.
evolve Opt-in model evolution (curate trajectories -> LoRA/DPO recipe).
evolve export Export a curated SFT or DPO dataset from trajectories.
evolve guard Watch evolution health; retract the most recent skill on a SIGNIFICANT regression (M19-A6).
evolve recipe Emit a runnable LoRA training recipe (train.py + README + requirements).
evolve refine GEPA-refine a skill from verified trajectories, gated on non-regressing transfer (M19-A5).
evolve rft One rejection-sampling fine-tuning round, gated by an honest A/B on two bench result files.
evolve status Show how much training signal the collected trajectories hold.
evolve tune Self-optimize the agent spec (OpenJarvis meta-search) against the daily scenarios.
explore Locate relevant code via the isolated Context Explorer subagent (FastContext-style).
features Show optional capabilities and what each needs (a key or a dependency).
find Search a repository by what code DOES, not by the string it contains.
fuse Run a prompt through the LLM-Fusion engine (panel -> judge -> synthesizer).
fusion-bench A/B the fusion engine: full vs selective (tokens + accuracy). Calls real models.
fusion-receipts Summarize persisted fusion receipts into an honest cost×quality curve.
guard Show the governance verdict (allow/warn/review/block) for an action.
hierarchy-bench Paired A/B: single-agent (all docs inline) vs the hierarchy (one worker per doc). Calls real models.
init First-run setup: create .env, set a provider key, and point you at a real example.
kanban Task board with worker lanes (backlog/doing/review/done).
kanban add Add a card to the backlog.
kanban board Show the board, column by column.
kanban learn Turn recurring tasks (from the experience buffer) into backlog cards.
kanban move Move a card to another column.
kanban rm Remove a card.
kanban run Dispatch backlog cards through their lanes (solve/crew). Requires a key.
lifecycle SDLC crew: plan -> build -> test -> review with verify-or-revert. Requires a key.
maturity Render the maturity scorecard: which coverage-IDs have their test file (presence, not passing).
mcp Configure MCP servers (persisted to .chimera/mcp.json). Terminal-first source of truth.
mcp add Add (or replace-by-name) an MCP server. Persists to .chimera/mcp.json — no connect.
mcp desktop Serve an MCP server on stdio that operates the RUNNING desktop app (for Claude Code/Desktop).
mcp list List configured MCP servers (name, command + args, env key names). No connect.
mcp remove Remove a configured MCP server by name.
mcp test Live-connect a configured server and print the tools it exposes (or a clear error).
measure Run the rulers this project measures itself with.
measure rag Recall@k of each retriever over a real folder — lexical, and vector when an embedder is set.
measure reranker Leave-one-out AUC of the success reranker — does it discriminate, or is it noise?
memory Curated long-term memory.
memory add Remember a fact (ADD / UPDATE / NOOP, deduped).
memory consolidate Merge clusters of similar memories into one LLM-summarised fact (opt-in write).
memory export Export all memory as JSON or Markdown, locally. Secrets are masked; metadata is left out.
memory graph Build an entity-relation graph from long-term memory and show it.
memory list List all memory items.
memory profile Show the consolidated cross-session user profile (persona facts).
memory prune Prune low-value memory under a budget. Dry-run by default; persona/profile facts are never pruned.
memory search Search memory (keyword).
memory-bench Measure recall@k as memory grows — lexical vs paraphrase.
memory-poison Ablate the memory-poisoning defenses: what reaches a LATER run's prompt, and unmarked.
meta Meta-agent: design a specialized agent blueprint for a task. Requires a key.
migrate Import config + skills from another agent; --apply also merges long-term memory.
models Model assignment: tier ladder (weak/mid/top), cost mode, and the multi-vendor catalog.
models catalog Browse the curated multi-vendor catalog (suggestions — any slug works).
models set Pin a tier to a model (or set the cost mode). Explicit pins always beat the mode.
orchestrate Hierarchical run: top model decomposes/synthesizes, budgeted mid workers execute.
pet Your virtual companion — a chimera that needs care.
pet feed Feed it (raises fullness).
pet new Adopt a fresh companion (resets stats).
pet play Play with it (raises happiness; costs energy + a little fullness).
pet rest Let it rest (restores energy).
pet status Check on your companion (stats drift while you're away).
playbook ACE strategy playbook — incremental, delta-curated guidance for the agent.
playbook add Manually add a bullet (a near-duplicate reinforces the existing one).
playbook curate Reflect on a run outcome and apply incremental deltas (add/reinforce/deprecate).
playbook refine Grow-and-refine: merge duplicate bullets and cap the size (deprecates the weakest).
playbook show Print the current active playbook (top strategies by score).
probe-select PROBE best-arm identification with a cheap-proxy control variate (M18-5).
profile Persistent user profile — the assistant's stable, cacheable preamble.
profile forget Remove a stored fact.
profile set Add a profile fact (name replaces; the list kinds append with dedup).
profile show Show the stored profile and the exact preamble sessions will receive.
project Run a project start-to-finish against a Spec (drift = acceptance authority).
project approve Approve the initial plan (default) or a paused high-risk card, then continue.
project deny Reject a paused high-risk card (parks it for review, escalates to a human).
project run Continue running a paused/escalated project (re-attempts a soft rail-stop).
project start Create a project from a spec and run it until it aligns or a rail stops it.
project status Show a project's status and its board.
project step Run exactly one iteration (cron-able).
redteam Red-team the injection defenses: attack success rate with vs without them.
report Reports counted by code — from this home's own logs, or read with the GitHub CLI — no model call.
report pr-watch Pull request watch: failing checks and new comments on your open pull requests, and failed runs
report weekly Weekly review: spend, runs, approvals and failing jobs over the last 7 days.
review [experimental] Review a change: findings first, P0 to P3, from a model of another family.
rubric-grade Grade an answer against an authorable rubric — weighted criteria with a required-criterion veto.
run Run a single-shot Tier-1 completion (no fusion). Requires a provider key.
sandbox-bench State-based bench: grade the final workspace state + count harmful side effects.
scenarios Run the daily right-hand scenario suite through a real chat session (live). Requires a key.
schema-bench Measure tool-schema token cost, full vs compacted (advertise-time). No model calls.
secrets Keep provider keys in the OS vault instead of a file.
secrets list What the OS vault holds — names only, never values.
secrets rm Remove one credential from the OS vault.
secrets set Put one credential in the OS vault.
serve Run the messaging gateway on HTTP, Discord, Telegram, Slack or Signal. Requires a key.
sessions List the conversations chimera chat and chimera tui have saved, under <home>/sessions.
skillcard-bench A/B reasoning with vs without injected TRS skill cards. Calls real models.
skills List the built-in skills.
skills-approve Approve/reactivate a learned skill after review (activates retrieval).
skills-bundle-disable Switch a bundle off, keeping it on disk.
skills-bundle-enable Switch an installed bundle on, so the agent may use it.
skills-bundles List the skill bundles installed on this machine, and where each came from.
skills-catalog Browse the installable skills from the wider Agent Skills ecosystem.
skills-evolve Reflectively evolve a skill's prompt template against graded instances (GEPA).
skills-export Export a learned skill to the open SKILL.md format (portable to the agent-skills ecosystem).
skills-import Import a SKILL.md into the store. A file imported by path is held pending for review.
skills-install Download a skill bundle from its source repository into your skills directory.
skills-library Browse the curated skill cards that ship with Chimera.
skills-lifecycle Run the measured skill-lifecycle loop (M18-4): promote proven provisionals, demote regressions.
skills-pending List learned skills held for review (e.g. distilled during a tainted run).
skills-retire Propose retiring under-performing skills — review-gated, never a delete.
skills-stats Per-skill usage stats (uses, successes, win rate) + retirement candidates.
skills-uninstall Delete an installed skill bundle and its files.
solve Tier-2: autonomously solve a task with plan + verify-or-revert. Requires a key.
solve-batch Solve several tasks concurrently, each in its own git worktree (Tier-3 isolation).
swe-bench-compare Honest A/B over two SWE-bench Verified-Mini reports on the SAME instance ids.
tools List the built-in native tools.
transfer-gate Promote a learned change only if it helps its tuned slice AND doesn't regress a holdout.
tui Launch the full-screen TUI — your right-hand. Requires a key.
version Show the Chimera version.
workflow Run a declarative workflow — a designed loop — from a YAML file. Requires a key.

a2a-card

Print Chimera's A2A Agent Card JSON (serve it at /.well-known/agent.json).

chimera a2a-card
Option Default
--url The A2A endpoint URL to advertise. 'http://127.0.0.1:8765/a2a'

acp

Serve Chimera to an editor over the Agent Client Protocol (stdio).

The mirror of what a Code-screen turn with provider: claude does: there we drive somebody else's agent, here somebody else's editor drives ours. Point Zed, JetBrains or Neovim at chimera acp and the loop, the verifier and the receipt are available without installing a second tool.

Nothing on this path may write to stdout — it IS the protocol. A stray print corrupts the frame the editor is parsing, and the symptom is an editor that hangs rather than output in the wrong place. The banner goes to stderr for the same reason the MCP server's does.

chimera acp
Option Default
--workspace, -w Directory the agent works in. '.'
--model Override the model for this session.
--max-steps Tool-calling steps per turn. 30

agent

Run the ReAct agent loop with native tools. Requires a provider key.

chimera agent TASK
Argument
TASK The task for the agent to accomplish.
Option Default
--model, -m Override the model slug.
--max-steps Max tool-calling steps. 8
--workspace, -w Workspace root for tools. '.'
--fuse Route deep-reasoning turns through fusion.
--guard Gate tool calls through the governance kernel.
--allow-tools Per-session allowlist: only these tools (comma-separated).
--deny-tools Per-session denylist: drop these tools (comma-separated).

agents

The agents you dispatch work to — as distinct from the one you converse with.

chimera agents

agents list

Show the registry.

chimera agents list

agents rm

Forget an agent. Cards already filed under its lane are left exactly where they are.

chimera agents rm AGENT_ID
Argument
AGENT_ID The agent to forget.

agents set

Add an agent, or replace the one with this id.

Replace rather than merge, matching the API: a partial write that kept what you left out would make clearing a pinned model impossible.

chimera agents set AGENT_ID
Argument
AGENT_ID Its handle: a lowercase slug. Also its Kanban lane.
Option Default
--name What to call it on screen. ''
--instructions Its role, in your words. ''
--model Pin a model; empty inherits the ladder. ''
--tools Comma-separated allowlist; empty means NO restriction. ''

app

Run the Chimera Desktop app: the HTTP+SSE API + the built React UI (needs the 'desktop' extra).

Serves the same-origin SPA (apps/desktop/dist) and a streaming chat API over the real agent stack. Install with pip install 'chimera-agent[desktop]' and build the UI once with npm --prefix apps/desktop run build.

chimera app
Option Default
--host Bind host (localhost by default). '127.0.0.1'
--port Bind port (0 = any free port; a busy port falls back to free). 8765
--model, -m Override the model slug.
--max-steps Max tool-calling steps per message. 6
--workspace, -w Workspace root for tools. Falls back to $CHIMERA_WORKSPACE, then the current directory — which is the wrong root for a packaged app, whose current directory is wherever its shortcut points.
--fuse Route turns through fusion (no token streaming).
--no-memory Don't recall long-term memory.
--cron Fire scheduled jobs while the app is open (proactivity). Default: the CHIMERA_APP_CRON setting (on). --no-cron makes the app purely reactive.
--open Open the app in your browser. True
--emit-port-file Write the final http://host:port URL to this file once bound (for a parent/sidecar).

approve

Answer a decision the agent is waiting on, from anywhere.

Without a terminal the approval gate collapsed to a refusal: ask degraded to deny, so every REVIEW verdict on the VPS, in a container or under cron was a no, and the mandate that says "confirm before billing, before a destructive migration, before touching RLS" had nothing to confirm with. The question is written down and sent to wherever this deployment delivers; this is how it gets answered.

Silence is still a refusal — a question times out. That is deliberate: a gate that reads silence as consent produces a record of an approval nobody gave.

chimera approve [REQUEST_ID]
Argument
REQUEST_ID The id from the message. Omit to list what is waiting.
Option Default
--yes, -y Approve it.
--no, -n Refuse it.
--show Print the whole question — the full action — and answer nothing.

assist

Your daily-driver assistant: cheap by default, escalates when it must.

Assist = chat with the second-brain defaults ON: the tier cascade routes chit-chat to cheap models and escalates hard asks; your persistent profile (chimera profile) is the stable preamble; memory, nudges and end-of-session consolidation are active. On exit it prints the session cost receipt — tier distribution + measured tokens — so 'cheap by default' is a number.

--max-usd bounds the whole run rather than one turn. Naming a model — --model or /model — pins it and turns the ladder off for as long as it is pinned, because the ladder is what chooses a model; it used to accept the slug and ignore it. /solve hands a task to the verified loop, inside the same ceiling.

chimera assist
Option Default
--model, -m Override the model slug.
--max-steps Max tool-calling steps per message. 6
--workspace, -w Workspace root for tools. '.'
--no-memory Don't recall long-term memory.
--no-cascade Disable tiered routing (single default model instead).
--max-usd Stop once this conversation has spent this much (the whole run, not one turn).
--write-region Comma-separated globs the file-writers may touch (e.g. 'src/**,*.py'). A write outside is refused — blocks an injected instruction from rewriting an unrelated file.

bench

Run the continuous-evolution benchmark on a demo task set. Requires a key.

chimera bench
Option Default
--limit Limit number of demo tasks (0 = all). 0
--model, -m Override the model slug.
--fuse Use the fusion engine as the solver.
--chain Run the stateful chained benchmark (error propagation).
--hard Use the hard suite (traps / propagating chain).
--rounds Re-run the suite N times; report stagnation + cost trend across rounds. 1

bench-compare

Report the honest A/B delta (+95% CI) between two benchmark result files.

Feed it the pass/fail from two runs on the SAME task IDs (e.g. a terminal-bench free-model baseline vs the same model driven by Chimera). Prints each arm's Wilson-bounded pass rate, the delta, its Newcombe CI, and whether the difference is significant. This is the number that proves (or doesn't) that the scaffolding lifts a weak model.

With --paired, the two lists are treated as aligned pairs (each index is one task replayed from an identical forked checkpoint), and the tighter McNemar/Wilson interval is reported — the payoff of running both arms from the same forked state.

chimera bench-compare BASELINE TREATMENT
Argument
BASELINE JSON file of the baseline arm's per-task pass/fail (list of bools, or {task: bool}).
TREATMENT JSON file of the treatment arm's per-task pass/fail.
Option Default
--baseline-name Label for the baseline arm. 'baseline'
--treatment-name Label for the treatment arm. 'chimera'
--paired Paired (McNemar) test: item i in both files is the SAME task replayed from an identical forked state — a tighter CI.

brief

Morning brief: parallel topic research through the hierarchy, one synthesized digest.

The recipe IS the decomposition (no top-model decompose call). Delegation receipts land in /delegations.jsonl — chimera delegations shows what the brief cost vs the inline counterfactual, measured.

chimera brief
Option Default
--recipe Brief recipe (YAML with topics). 'examples/morning_brief/brief.yaml'
--out Write the digest to this file (default: print only).
--max-workers Parallel research workers. 4

cascade-bench

Four-arm bench: weak-only vs mid-only vs cascade vs fusion. Calls real models.

Published criterion (stated up front): cascade >= mid-only pass rate at materially lower tokens-per-pass. The number reported is whatever is measured.

chimera cascade-bench
Option Default
--tasks Task suite: hard demo.

chat

Interactive multi-turn chat — your terminal right-hand. Requires a key.

The whole conversation is saved after every turn, under <home>/sessions, and picked up again on the next run. chimera sessions lists the threads and chimera chat -s <id> resumes one; GET /api/sessions serves the same files to any HTTP client.

Coding conversations in the desktop app are a different store — <home>/code_sessions, which keeps the model's own message list and its turn receipts rather than prose pairs — so a thread does not travel between the two.

A resumed turn is labelled as restored in the next prompt, and one that ran while untrusted content was in the conversation comes back inside the data fence; a turn saved before that was recorded is treated the same way, because nothing measured it — and a dim line under each reply now says so on screen. Memory recall is scoped to --workspace: that folder's facts, plus the ones stored with no project at all.

--max-usd bounds the whole thread rather than one turn, and /solve hands the conversation's task to the same verified loop chimera solve runs, inside that same ceiling.

chimera chat
Option Default
--model, -m Override the model slug.
--max-steps Max tool-calling steps per message. 6
--workspace, -w Workspace root for tools. '.'
--fuse Route deep-reasoning turns through fusion.
--cascade Tiered routing: weak -> gate -> mid -> gate -> fusion (cheap by default).
--no-memory Don't recall long-term memory.
--session, -s Resume a specific session id (see 'chimera sessions').
--new Start a fresh session instead of resuming.
--max-usd Stop once this conversation has spent this much (the whole thread, not one turn).
--write-region Comma-separated globs the file-writers may touch (e.g. 'src/**,*.py'). A write outside is refused — blocks an injected instruction from rewriting an unrelated file.

code

Continue a desktop Code conversation from this terminal, through the running app.

chimera code

code list

The desktop app's Code conversations, newest first, with the id resume takes.

chimera code list
Option Default
--limit, -n How many conversations, newest first. 20

code resume

Continue a desktop Code conversation here; the app runs the turn and shows it too.

Needs the app open with Settings > "Allow Claude to operate this app" on. Approval questions are printed; answering them here also needs "Full control", otherwise answer them in the app. The turns run on the models configured in the app: the bridge this command speaks through takes no model choice.

chimera code resume SESSION_ID
Argument
SESSION_ID The conversation's id (or a prefix only it has).
Option Default
--message, -m Send this one message and exit. Omit to keep talking.

context-curve

Did runs carrying more context do worse? Measured on THIS machine's own logs.

Answers with "not enough data" until the pre-registered floors are met — see bench/context_curve/PREREGISTRATION.md, which fixed those floors before any data existed.

chimera context-curve
Option Default
--traces Path to traces.jsonl (default: CHIMERA_HOME).
--runs Path to runs.jsonl (default: CHIMERA_HOME).
--json Print the raw result instead of a table.

crew

Run a multi-agent crew on a task (Tier 3). Requires a provider key.

chimera crew TASK
Argument
TASK The task for the crew.
Option Default
--mode sequential supervisor
--model, -m Override the model slug.
--fuse Use the fusion engine as the backend.

crew-isolated

Tier-3: tool-using workers attempt ONE task, each in its own git worktree, verify-gated.

Define workers with repeated --worker 'name:instruction'. Each runs a real agent loop (search/read/edit) against an isolated checkout; a worker whose check fails is rejected (its edits discarded). Of the workers that pass, ONE lands whole — a check that ran beats one that could not, then the smallest diff, then the order given — and --verify runs again on the merged workspace. With --merge-all, every approved worker's files land instead, and files two of them changed are flagged as conflicts. Needs a git repo to isolate.

chimera crew-isolated TASK
Argument
TASK The task every worker attempts (or divides, with --merge-all).
Option Default
--worker, -W A worker as 'name:instruction'; repeatable. Each edits in its own worktree.
--workspace, -w Repository root (a git repo, to isolate). '.'
--model, -m Override the model slug.
--verify Per-worker gate: shell command run in each worktree (exit 0 to merge).
--max-steps Max tool-calling steps per worker. 6
--max-workers Max concurrent isolated workers. 4
--synthesize A supervisor folds the merged results into one unified report.
--merge-all Workers do DISJOINT parts: land every approved worker's files instead of one worker's whole tree (files two of them changed land from neither).
--fuse Route worker turns through fusion.
--taint Arm each worker's adaptive allowlist (dangerous-when-tainted tools require approval). The cross-agent collusion monitor runs regardless — it's always on for fan-out.

cron

Manage scheduled jobs (crons and event SOPs).

chimera cron

cron add

Add a cron, event- or webhook-triggered job.

--verify is what turns a scheduled job into a run the harness governs. CronJob has carried the field since the harness landed and nothing could write it — not this command, not the HTTP route — so for every user the gate was permanently unarmed.

chimera cron add NAME SCHEDULE ACTION
Argument
NAME A human-readable name.
SCHEDULE Cron expression, or an event/webhook name.
ACTION What to do (task description / skill).
Option Default
--event Treat SCHEDULE as an event name.
--webhook Fire on POST /webhook/ (needs 'chimera serve').
--verify Gate: shell command run in the job's folder after the dispatch (exit 0 to keep the work, non-zero to revert it). Empty = no gate, which is the previous behaviour. ''
--max-attempts Attempts per dispatch. Worth raising only with --verify: without a gate nothing can tell a failed attempt from a finished one. 1
--notify When the answer is posted to the job's destination: always (every answer except the job's own 'nothing new' reply), on_change (skip an answer identical to the last one delivered), or failures_only. The result file gets every answer either way. Not with --webhook: a webhook job answers through the chat gateway. 'always'
--tools Comma-separated tools this job may use; the rest are removed from its registry. Omit for every tool (the previous behaviour). Refused with --webhook: a webhook job runs through the chat gateway, which does not apply the list.
--deliver-to Chat webhook URL (Discord or Slack) the job's answers are posted to, per --notify; a run that could not run or finish is announced there too. The URL is a credential and is never printed in full. Refused with --webhook: that job answers through the chat gateway.

cron disable

Disable a job without deleting it.

chimera cron disable JOB_ID
Argument
JOB_ID The job id to disable.

cron doctor

Ask the schedule what it is not telling you: what never ran, and what ran and lost.

Every other honesty mechanism here sits downstream of a run having happened. This is the one question about the run that did not — and about the one that happens on time, forever, and fails every time, which looks healthier than the first from any field that existed before.

It is a question, not a watcher: nothing notices while this process is down, for the same reason a crashed process cannot log its own crash. What it gives you is an honest answer the moment you ask.

chimera cron doctor
Option Default
--grace How late a job may be before it counts as missed. 10.0
--check Exit 1 when a job is late or failing, so a watcher outside Chimera alerts only then.

cron enable

Enable a job (e.g. an agent-proposed one) and schedule its next run.

chimera cron enable JOB_ID
Argument
JOB_ID The job id to enable.

cron fire

Run every job registered for an event.

Event jobs had no dispatcher. cron add --event deploy accepted the job and cron list showed it enabled, but nothing in the package ever called fire_event — so the job simply never ran, and its silence was indistinguishable from that of a job whose time had not come. The cron trigger has the daemon and the webhook trigger has the webhook server; this is the third one's.

Meant to be called from wherever the event actually happens — a git hook, a deploy step, a CI job. Dispatch is the same one the daemon uses, so a fired job behaves exactly like a scheduled one: same agent, same spend caps, same receipt.

chimera cron fire EVENT
Argument
EVENT The event name to fire (as given to cron add --event).
Option Default
--model, -m Model for the dispatched jobs.
--max-steps Max tool-calling steps per job. 6
--workspace, -w Workspace root for tools. '.'

cron kill

Stop a job's running (or next) dispatch — one run, not the schedule.

disable takes the job off the clock; kill answers the other question: the job is running RIGHT NOW and must stop. The daemon's worker polls the flag between steps, the dispatch it stops deletes it, and the run ends cancelled — which counts as neither a failure nor a success, so a kill cannot ride the failure counter into the brake.

chimera cron kill JOB_ID
Argument
JOB_ID The job id to stop.

cron learn

Propose crons from recurring tasks and create the ones you confirm.

Each proposal is shown for explicit confirmation (the human-in-the-loop approval that keeps automation creation under control); confirmed jobs are validated and created enabled. --yes confirms all (use deliberately).

chimera cron learn
Option Default
--min Min repeats to propose. 3
--schedule Override the suggested cron schedule.
--yes, -y Create every proposal without prompting.

cron list

List scheduled jobs.

chimera cron list

cron remove

Remove a scheduled job by id.

chimera cron remove JOB_ID
Argument
JOB_ID The job id to remove.

decide

Ask typed questions — yes/no, a choice, a score — and get probabilities back.

Every question is read on its own, decision-first; a question the linter rejects is refused before any call. noul is P(yes); confidence describes how peaked the probabilities are and is not a probability of being right. A number is calibrated only where a map exists for exactly this question.

Exit codes: 0 every question answered; 1 at least one question failed (a halt: the backend was down or the state overflowed) — the output is still printed in full first; 2 usage, or a question the linter refuses before any call.

Measured on a ruler we did not build (bench/jevbench_local, the 231 public JevBench items): the default local backend answers 0.619 of them right (Jev 1.13: 0.866), 0.324 on the hard tier, with a raw top-label ECE of 0.218. Options that share a first token cannot be read locally: name them so their first words differ.

chimera decide
Option Default
--questions, -q JSON file: the questions (request shape, no state).
--state, -s The state to ask about. ''
--jsonl Ask about every line of this JSONL file instead.
--field With --jsonl: the field that holds the state. 'state'
--out, -o With --jsonl: write results here (default: stdout).
--decision Name the decision (the key a calibration map is found by). ''
--no-log Do not write the answers to the decision log.

decisions

Typed decisions: which model answers them, the log of what they answered, labels, a report and a refit.

chimera decisions

decisions label

Say what was true for one answer. A later label for the same id replaces an earlier one.

chimera decisions label ENTRY_ID
Argument
ENTRY_ID The id from chimera decisions log or the approval card.
Option Default
--yes, -y The question's event happened (governance: it WAS dangerous).
--no, -n It did not (governance: it was NOT dangerous).
--note Why, for whoever reads the label later. ''

decisions log

The latest answers, newest last, with their label when one was given.

chimera decisions log
Option Default
--limit, -n How many of the latest answers. 20
--unlabelled Only answers without a label.

decisions models

The System One models OpenRouter lists, which one is active, and which carry a calibration map.

chimera decisions models

decisions refit

Fit this deployment's own map on its labelled answers — pooled with the shipped rows while it has fewer than 20 labels of a class. Prints before writing; writes only with --write.

chimera decisions refit
Option Default
--write Save the refitted maps to /decisions/maps.json.

decisions report

What the log holds: availability, the review budget, label coverage, and — where labels exist — catch and false refusal at the REVIEW threshold.

chimera decisions report

decisions use

Choose the backend (and model) that answers typed decisions — written to .env in this folder, the same pair and the same check as the desktop's System One card.

chimera decisions use BACKEND [MODEL]
Argument
BACKEND local_logprob
MODEL Empty = the backend's measured default. For openrouter_decisions, a listed slug.

delegations

Measured vs counterfactual across delegations — what the hierarchy actually saved.

chimera delegations
Option Default
--path Receipts file (default: /delegations.jsonl).

deliver

Deliverable Mode: produce a polished, self-contained artifact. Requires a key.

chimera deliver REQUEST
Argument
REQUEST What to produce (a report, plan, spec, README...).
Option Default
--out, -o Write the deliverable to this file.
--format, -f md txt
--model, -m Override the model slug.
--fuse Use the fusion engine for higher quality.

doctor

Check the environment and configuration. With --fix, repair safe setup issues.

--probe is off by default and that is deliberate: doctor should stay instant, offline and free. What it buys when you ask for it is the difference between a claim and a measurement — "Ready" below is an assertion about the NAME of an environment variable, so a revoked key, an account with no credit, or a value pasted with a trailing space all pass it and fail on the first real call. The argument for measuring is already written in config_api.pricing_capability a few files over: the time to find out is while reading the doctor, not when a 3 a.m. cron stalls.

chimera doctor
Option Default
--fix Auto-repair safe setup issues (state dir, .env scaffold).
--probe Actually call the provider once, instead of trusting the key's name.

drift

Drift gate: check the workspace against a spec (Spec Growth). Exit 1 on drift.

chimera drift SPEC
Argument
SPEC Spec YAML file.
Option Default
--workspace, -w Workspace root. '.'
--only Check only this requirement id (project cards).

evoclaw

Stress-test continuous-evolution degradation: naive vs guarded. Requires a key.

chimera evoclaw
Option Default
--length Number of chained steps. 12
--model, -m Override the model slug.
--retries Verify-or-revert retries per step (guarded). 2

evolve

Opt-in model evolution (curate trajectories -> LoRA/DPO recipe).

chimera evolve

evolve export

Export a curated SFT or DPO dataset from trajectories.

chimera evolve export
Option Default
--out Output JSONL path.
--format sft dpo
--traj Trajectory JSONL (default: /trajectories.jsonl).
--min-reward Drop examples below this reward. 0.0
--no-dedup Keep duplicate examples.
--min-margin DPO: min reward margin chosen − rejected. 0.0
--min-steps Recipe: keep only traces with >= N steps. 0
--diverse Recipe: at most one SFT example per task.
--min-process Keep only traces whose step-following score >= this (SkillCoach).

evolve guard

Watch evolution health; retract the most recent skill on a SIGNIFICANT regression (M19-A6).

chimera evolve guard
Option Default
--limit Limit demo tasks (0 = all). 0
--model, -m Override the model slug.
--cost-drift-tol Also roll back if second-half mean cost exceeds first by more than this.
--apply Retire the most recent skill IF a SIGNIFICANT regression is measured.

evolve recipe

Emit a runnable LoRA training recipe (train.py + README + requirements).

chimera evolve recipe
Option Default
--out Directory for the training recipe.
--format sft dpo
--base-model 'meta-llama/Llama-3.1-8B-Instruct'
--dataset Dataset filename the script reads. 'dataset.jsonl'

evolve refine

GEPA-refine a skill from verified trajectories, gated on non-regressing transfer (M19-A5).

chimera evolve refine SKILL_NAME
Argument
SKILL_NAME Name of the learned skill to refine.
Option Default
--traj Trajectory JSONL (default: /trajectories.jsonl).
--model, -m Override the model.
--budget GEPA rollout budget. 20
--min-reward Only mine trajectories at/above this reward (1.0 = verified). 1.0
--apply Persist the refined skill IF it passes the transfer gate.

evolve rft

One rejection-sampling fine-tuning round, gated by an honest A/B on two bench result files.

Rejection-samples the collected trajectories (successes at/above the reward bar), then promotes the round ONLY if the candidate beats the baseline with a confidence interval that excludes zero — no lift, no promotion, no training on noise. Artifacts are withheld for an unpromoted round unless --force. Feed --baseline/--candidate the pass/fail lists two bench runs produce.

chimera evolve rft
Option Default
--baseline JSON list of baseline bench pass/fail.
--candidate JSON list of candidate bench pass/fail.
--traj Trajectory JSONL (default: /trajectories.jsonl).
--min-reward Rejection-sampling reward bar. 0.5
--min-examples Accepted examples needed to gate. 30
--top-k Keep at most this many accepted per prompt (0 = all). 0
--out If promoted, write dataset + recipe here.
--force Export even if the round is not promoted.

evolve status

Show how much training signal the collected trajectories hold.

chimera evolve status
Option Default
--traj Trajectory JSONL (default: /trajectories.jsonl).
--min-reward Drop examples below this reward. 0.0
--min-examples Examples needed before training is worth it. 30

evolve tune

Self-optimize the agent spec (OpenJarvis meta-search) against the daily scenarios.

Each round a model proposes a coordinated edit to the spec; the candidate is scored on the daily scenarios and kept only on non-regression. Uses real model calls.

chimera evolve tune
Option Default
--rounds Meta-search rounds. 2
--model Base model for the spec.
--max-steps Initial runtime step budget. 8
--k Suite runs per candidate — one samples, two alert, three decide. Multiplies cost. 3

explore

Locate relevant code via the isolated Context Explorer subagent (FastContext-style).

Returns only a compact file:line evidence block — the exploration turns never touch your context. A cheap model is usually the right call here; localization is a narrow task. With CHIMERA_EXPLORER_CONTRACT on, it returns findings with a location each, a gaps section, and a check of every cited location against the workspace.

chimera explore QUERY
Argument
QUERY What to locate in the repository.
Option Default
--workspace, -w Repository root to explore. '.'
--model, -m Model for the explorer (a cheap one is fine).
--max-turns Max exploration turns. 8
--thoroughness quick, medium or thorough: halves, keeps or doubles --max-turns. Read only when CHIMERA_EXPLORER_CONTRACT is on. 'medium'

features

Show optional capabilities and what each needs (a key or a dependency).

chimera features

find

Search a repository by what code DOES, not by the string it contains.

chimera/rag/ has been in the tree since 0.44.0 — symbol-level chunking over Python's AST, one SQLite file with an FTS5 index, RRF fusion — measured, documented, and reachable from nothing. A library with no entrance is a library nobody has. This is the entrance.

Keyword by default; --semantic fuses it with embeddings, and the fusion is what was measured and adopted in bench/rag/RESULTS.md: hybrid 0.5050 against keyword 0.4425 on this repository, +6.25 pp paired over 400 probes, McNemar p = 1.7e-04.

--semantic means HYBRID, never vectors alone, and that is the measurement rather than a preference: the vector arm on its own scored 0.4100 — worse than keyword. Every point of the win comes from fusing two rankings that are wrong about different things. A flag that gave you the vector arm would be a flag that made your search worse.

The recall figure is printed with every search because it is per-corpus and per-embedder: the same harness measures 0.4750 on this repository as it stood three weeks ago, and there is no conversion from one embedding model's vector space to another's.

chimera find QUERY
Argument
QUERY What you are looking for, in words.
Option Default
--path Repository to search. '.'
--k How many results. 8
--reindex Rebuild the index before searching.
--semantic Fuse keyword with embeddings. Costs money to index.

fuse

Run a prompt through the LLM-Fusion engine (panel -> judge -> synthesizer).

chimera fuse PROMPT
Argument
PROMPT The prompt to run through the fusion engine.
Option Default
--show-panel Show panel answers + judge analysis.
--selective Override selective fusion (default: from settings).
--best-of Cheap fusion: sample ONE model N times and take the consensus (self-consistency), instead of a multi-model panel. 1
--verify-select With --best-of: pick the best sample by a verifier score instead of majority vote (Weaver-lite).
--model, -m Model for --best-of self-consistency.
--show-cost Print the itemized receipt: per-advisor cost at each model's rate.
--receipt Append the run's cost receipt to this JSONL (for cost×quality analysis).

fusion-bench

A/B the fusion engine: full vs selective (tokens + accuracy). Calls real models.

chimera fusion-bench
Option Default
--tasks Task suite: hard demo.

fusion-receipts

Summarize persisted fusion receipts into an honest cost×quality curve.

chimera fusion-receipts PATH
Argument
PATH JSONL of receipts written by fuse --receipt.

guard

Show the governance verdict (allow/warn/review/block) for an action.

chimera guard ACTION
Argument
ACTION The action/command to evaluate.

hierarchy-bench

Paired A/B: single-agent (all docs inline) vs the hierarchy (one worker per doc). Calls real models.

Both arms run on the SAME model so the comparison isolates the ORCHESTRATION (minimal-context scoping + budgets + contracts), not model strength. Quality = paired McNemar/Wilson (the only place "significant" appears); tokens = measured totals per arm, with no significance claim on cost.

--multistep switches to the companion suite where the token crossover lives: a single agent re-sends every document on every turn (cost grows with turns), while scoped workers pay each doc ~once — and prices the measured cache reduction via the caching model.

chimera hierarchy-bench
Option Default
--model, -m Mid/worker model — BOTH arms use it, to isolate orchestration. Defaults to the tier ladder's mid.
--top-model Top model for synthesis. Defaults to --model (same family keeps the isolation).
--tasks Comma-separated task ids to filter (default: all 10 synthetic tasks). ''
--max-workers Max concurrent workers in the hierarchy arm. 4
--out Write the JSON summary to this path.
--multistep Run the MULTI-STEP suite instead (single growing context vs per-step scoped workers, over large docs) — the regime where the hierarchy actually saves tokens. Also reports a caching-aware dollar reduction.

init

First-run setup: create .env, set a provider key, and point you at a real example.

chimera init
Option Default
--provider Which provider the key is for (openrouter, openai, ...). 'openrouter'
--key API key for --provider.
--openrouter-key Your OpenRouter API key (same as --provider openrouter --key).
--model Default model slug to set (optional).
--yes, -y Non-interactive: never prompt.
--home Project dir for the .env (default: cwd).

kanban

Task board with worker lanes (backlog/doing/review/done).

chimera kanban

kanban add

Add a card to the backlog.

chimera kanban add TITLE
Argument
TITLE Short card title.
Option Default
--action, -a Task text to run (defaults to title).
--lane, -l Who works it: solve crew
--verify Verify command for the solve lane (exit 0).

kanban board

Show the board, column by column.

chimera kanban board

kanban learn

Turn recurring tasks (from the experience buffer) into backlog cards.

Uses the cron-learner's recurrence detector; each card is confirmed (or --yes), and a task already on the board is skipped — so re-running is safe.

chimera kanban learn
Option Default
--min Min repeats to turn into a card. 3
--lane, -l Lane for the created cards. 'solve'
--yes, -y Add every card without prompting.

kanban move

Move a card to another column.

chimera kanban move CARD_ID COLUMN
Argument
CARD_ID Card id.
COLUMN backlog

kanban rm

Remove a card.

chimera kanban rm CARD_ID
Argument
CARD_ID Card id.

kanban run

Dispatch backlog cards through their lanes (solve/crew). Requires a key.

chimera kanban run
Option Default
--limit, -n Max backlog cards to dispatch.
--workspace, -w Workspace for the solve lane. '.'
--model, -m Override the model slug.
--workers, -j Work this many cards at once, each in its own git worktree. 1

lifecycle

SDLC crew: plan -> build -> test -> review with verify-or-revert. Requires a key.

chimera lifecycle TASK
Argument
TASK The feature/task to take through the SDLC.
Option Default
--verify Test command for the test stage (exit 0).
--workspace, -w Workspace root. '.'
--model, -m Override the model slug.
--max-attempts Build/test verify-or-revert budget. 2

maturity

Render the maturity scorecard: which coverage-IDs have their test file (presence, not passing).

chimera maturity
Option Default
--tests Path to the tests directory (the evidence base). 'tests'

mcp

Configure MCP servers (persisted to .chimera/mcp.json). Terminal-first source of truth.

chimera mcp

mcp add

Add (or replace-by-name) an MCP server. Persists to .chimera/mcp.json — no connect.

chimera mcp add NAME
Argument
NAME A unique name for the server (namespaces its tools).
Option Default
--command, -c The launch command (e.g. npx, uvx, python).
--arg, -a A command argument (repeatable).
--env, -e An env var as K=V (repeatable).

mcp desktop

Serve an MCP server on stdio that operates the RUNNING desktop app (for Claude Code/Desktop).

Needs the app open with Settings > "Allow Claude to operate this app" on; the tools then call the app's local bridge. Register it with: claude mcp add chimera-desktop -- chimera mcp desktop

chimera mcp desktop

mcp list

List configured MCP servers (name, command + args, env key names). No connect.

chimera mcp list

mcp remove

Remove a configured MCP server by name.

chimera mcp remove NAME
Argument
NAME The server name to remove.

mcp test

Live-connect a configured server and print the tools it exposes (or a clear error).

This is the ONLY MCP subcommand that connects (spawns the server + runs the async handshake). It is the sole honest proof a server is reachable. Needs the 'mcp' extra and the server's own runtime (e.g. Node for an npx server).

chimera mcp test NAME
Argument
NAME The configured server to live-test.
Option Default
--timeout Connect timeout in seconds. 12.0

measure

Run the rulers this project measures itself with.

chimera measure

measure rag

Recall@k of each retriever over a real folder — lexical, and vector when an embedder is set.

This is the measurement chimera/rag/__init__.py names when it says the retriever's existence is not a claim that it helps. That sentence pointed at a module you could not run: rag_bench had no caller outside its own test and was not exported from chimera.eval.

Without --semantic no embedder is passed, so the vector and hybrid figures come back as None rather than zero — an embedder that was never called did not fail, and printing 0.0 invites the wrong conclusion.

With it, the run that bench/rag/RESULTS.md reports is reproducible from the CLI rather than from a script somebody has to write. It costs an embedding pass over the corpus: about two cents for this repository's 3,459 chunks and 400 probes, and the figure it produces belongs to the embedder that produced it — vector spaces do not convert between models.

chimera measure rag ROOT
Argument
ROOT Folder to index and probe.
Option Default
--k Retrieve this many chunks per probe. 10
--max-probes Cap the probe count; each one is a query. 200
--semantic Measure the vector and hybrid arms too. Costs money.

measure reranker

Leave-one-out AUC of the success reranker — does it discriminate, or is it noise?

chimera/evolution/reranker.py says to measure with this BEFORE putting the reranker in a hot path. It was prose pointing at an unreachable module.

AUC of 0.5 is a coin flip. A reranker at 0.5 is not a weak reranker, it is not a reranker.

chimera measure reranker CORPUS
Argument
CORPUS JSONL of {query, text, success} records.
Option Default
--k Rank cut-off for the leave-one-out scoring. 5

memory

Curated long-term memory.

chimera memory

memory add

Remember a fact (ADD / UPDATE / NOOP, deduped).

chimera memory add CONTENT
Argument
CONTENT The fact to remember.
Option Default
--key Optional dedup key.
--persona Store as a persona fact (part of the cross-session profile).

memory consolidate

Merge clusters of similar memories into one LLM-summarised fact (opt-in write).

chimera memory consolidate
Option Default
--threshold Similarity (Jaccard) to cluster facts; lower = merges more. 0.5
--dry-run Only list the clusters that would be merged (no model call, no write).

memory export

Export all memory as JSON or Markdown, locally. Secrets are masked; metadata is left out.

chimera memory export
Option Default
--format json markdown
--out Write to this file (default: print to stdout).

memory graph

Build an entity-relation graph from long-term memory and show it.

chimera memory graph
Option Default
--entity, -e Show relations for one entity.

memory list

List all memory items.

chimera memory list

memory profile

Show the consolidated cross-session user profile (persona facts).

chimera memory profile

memory prune

Prune low-value memory under a budget. Dry-run by default; persona/profile facts are never pruned.

chimera memory prune
Option Default
--max Keep the N highest-value memories. 50
--apply Actually delete. Default is a dry-run preview (no data lost).

Search memory (keyword).

chimera memory search QUERY
Argument
QUERY Search query.
Option Default
--k Max results. 5

memory-bench

Measure recall@k as memory grows — lexical vs paraphrase.

Default (keyword search, no key needed) surfaces the honest ceiling: exact-token recall holds at scale, but paraphrase recall collapses. Pass --semantic to re-run with the embedding recall path and watch the paraphrase column lift — that delta is the whole point of M11b.

chimera memory-bench
Option Default
--sizes Comma-separated memory sizes to sweep. '50,200,1000'
--semantic Use embedding recall (needs an embeddings key) to measure the lift.

memory-poison

Ablate the memory-poisoning defenses: what reaches a LATER run's prompt, and unmarked.

No key needed, nothing leaves the machine. redteam measures one run — content arrives untrusted, the harmful call is refused, and the whole picture ends with the process. This measures the other shape: run A stores what it "learned" from a poisoned page, run B asks an unrelated question days later, and recall hands the planted fact to the model.

The headline is what arrives unmarked, not what is blocked. A poisoned fact carrying its origin is one the model was warned about; an unlabelled one is indistinguishable from something the agent verified itself. Each of the three layers (taint / gate / label) is switched off in turn, because a single number would be compatible with any of them doing nothing.

See bench/memory_poison/PREREGISTRATION.md for the thresholds, fixed before the first run.

chimera memory-poison

meta

Meta-agent: design a specialized agent blueprint for a task. Requires a key.

chimera meta TASK
Argument
TASK The task to design a specialized agent for.
Option Default
--model, -m Override the model slug.

migrate

Import config + skills from another agent; --apply also merges long-term memory.

claude imports memory only (CLAUDE.md and memory/*.md), as unverified facts, never persona: the dry-run lists every fact it would write.

chimera migrate SOURCE PATH
Argument
SOURCE Source agent: hermes
PATH Path to the source agent's home directory (for claude: ~/.claude or a project).
Option Default
--apply Write artifacts (default: dry-run preview).
--home Target Chimera home (default: from config).

models

Model assignment: tier ladder (weak/mid/top), cost mode, and the multi-vendor catalog.

chimera models

models catalog

Browse the curated multi-vendor catalog (suggestions — any slug works).

chimera models catalog
Option Default
--tier Filter: weak, mid, or top.
--vendor Filter by vendor substring.

models set

Pin a tier to a model (or set the cost mode). Explicit pins always beat the mode.

chimera models set ROLE VALUE
Argument
ROLE weak
VALUE A model slug (any vendor), 'auto' to unpin, or a cost mode for 'mode'.

orchestrate

Hierarchical run: top model decomposes/synthesizes, budgeted mid workers execute.

Write-shaped and trivial tasks FALL BACK to the single-agent path by design (the evidence says multi-agent loses there); the fallback is logged with its counterfactual so chimera delegations shows the decision.

chimera orchestrate TASK
Argument
TASK The task (read-heavy multi-part tasks benefit most).
Option Default
--max-workers Parallel worker cap. 4
--budget Token budget per delegation (default: settings).
--dry-run Show classification + decomposition + estimate; zero worker spend.
--verify-model Model slug for the spot-check auditor (a DISTINCT/cross-provider model that grades a worker's summary against its raw output). Default: the weak tier.

pet

Your virtual companion — a chimera that needs care.

chimera pet

pet feed

Feed it (raises fullness).

chimera pet feed

pet new

Adopt a fresh companion (resets stats).

chimera pet new
Option Default
--name Companion name. 'Chimi'
--species Companion species. 'chimera'

pet play

Play with it (raises happiness; costs energy + a little fullness).

chimera pet play

pet rest

Let it rest (restores energy).

chimera pet rest

pet status

Check on your companion (stats drift while you're away).

chimera pet status

playbook

ACE strategy playbook — incremental, delta-curated guidance for the agent.

chimera playbook

playbook add

Manually add a bullet (a near-duplicate reinforces the existing one).

chimera playbook add CONTENT
Argument
CONTENT The strategy/pitfall bullet to add.
Option Default
--section strategy pitfall

playbook curate

Reflect on a run outcome and apply incremental deltas (add/reinforce/deprecate).

chimera playbook curate
Option Default
--task The task the outcome is for.
--outcome What happened (success/failure + details).
--model Model slug for the reflect+curate call.

playbook refine

Grow-and-refine: merge duplicate bullets and cap the size (deprecates the weakest).

chimera playbook refine

playbook show

Print the current active playbook (top strategies by score).

chimera playbook show

probe-select

PROBE best-arm identification with a cheap-proxy control variate (M18-5).

"Which model/config is best?" where each expensive reward (a real grade) is paired with a cheap proxy (a weak judge) of unknown correlation. PROBE uses the proxy as a control variate so the estimate needs FEWER expensive draws the better the proxy correlates — and stays unbiased when the proxy is useless. Prints each arm's adjusted mean ± interval, the winner, and — if not yet confident — the arm to sample next. Feed it recorded (proxy, reward) observations from a bench.

chimera probe-select [DATA]
Argument
DATA JSON: {"arm": [[proxy, reward-or-null], ...], ...}. Omit when using --from-log.
Option Default
--from-log Read observations from a ProbeLog JSONL (e.g. /probe.jsonl written by solve --probe-log).
--delta Confidence level (smaller = stricter). 0.1
--min-reward Expensive rewards required per arm before deciding. 2

profile

Persistent user profile — the assistant's stable, cacheable preamble.

chimera profile

profile forget

Remove a stored fact.

chimera profile forget VALUE
Argument
VALUE The exact fact to remove (or 'name' to clear the name).

profile set

Add a profile fact (name replaces; the list kinds append with dedup).

chimera profile set KIND VALUE
Argument
KIND name
VALUE The fact to store.

profile show

Show the stored profile and the exact preamble sessions will receive.

chimera profile show

project

Run a project start-to-finish against a Spec (drift = acceptance authority).

chimera project

project approve

Approve the initial plan (default) or a paused high-risk card, then continue.

chimera project approve PROJECT_ID
Argument
PROJECT_ID Project id.
Option Default
--card Approve a specific high-risk card instead of the plan.
--model, -m Override the solve model.

project deny

Reject a paused high-risk card (parks it for review, escalates to a human).

chimera project deny PROJECT_ID
Argument
PROJECT_ID Project id.
Option Default
--card The high-risk card to reject.

project run

Continue running a paused/escalated project (re-attempts a soft rail-stop).

chimera project run PROJECT_ID
Argument
PROJECT_ID Project id.
Option Default
--model, -m Override the solve model.

project start

Create a project from a spec and run it until it aligns or a rail stops it.

chimera project start SPEC
Argument
SPEC Spec YAML (the acceptance authority).
Option Default
--workspace, -w Project workspace root. '.'
--model, -m Override the solve model.
--max-iterations Hard rail on card runs. 20
--yes, -y Skip the initial plan-approval pause (auto-approve).

project status

Show a project's status and its board.

chimera project status PROJECT_ID
Argument
PROJECT_ID Project id.

project step

Run exactly one iteration (cron-able).

chimera project step PROJECT_ID
Argument
PROJECT_ID Project id.
Option Default
--model, -m Override the solve model.

redteam

Red-team the injection defenses: attack success rate with vs without them.

No key needed — measures whether the governance layer blocks a harmful tool call once a run is tainted (defense-in-depth coverage), not model susceptibility.

chimera redteam

report

Reports counted by code — from this home's own logs, or read with the GitHub CLI — no model call.

chimera report

report pr-watch

Pull request watch: failing checks and new comments on your open pull requests, and failed runs on the default branch — read with the GitHub CLI, never acted on.

Without --print/--json this registers the watch for WORKSPACE as an hourly job, DISABLED, once per repository: it runs only after chimera cron enable <id>, posts only where --deliver-to says, and posts again only when the summary changes. Nothing is pushed, commented, merged or re-run, and other people's comments are quoted inside the data fence, as data.

chimera report pr-watch
Option Default
--workspace, -w The repository to watch (a folder inside a git checkout whose origin is on GitHub). '.'
--print Look now and print the summary instead of proposing the job. Reads only; remembers nothing.
--json Look now and print what was found as JSON. Reads only; remembers nothing.
--deliver-to Chat webhook URL (Discord or Slack) the job posts to. Stored on the proposal; never printed in full.
--lang pt or en. Default: the owner's identity language (Portuguese unless it names another).

report weekly

Weekly review: spend, runs, approvals and failing jobs over the last 7 days.

Every number is computed by code from the same logs the app's screens read (usage.jsonl, runs.jsonl, approvals/history.jsonl, scheduler/jobs.json); no model writes or restates any of them. Without --print this registers the weekly job — Mondays 09:00, DISABLED — once: it runs only after chimera cron enable <id>, and posts only where --deliver-to says.

chimera report weekly
Option Default
--print Print the last 7 days' review now instead of proposing the weekly job. Reads only.
--deliver-to Chat webhook URL (Discord or Slack) the weekly job posts to. Stored on the proposal; never printed in full.
--lang pt or en. Default: the owner's identity language (Portuguese unless it names another).

review

[experimental] Review a change: findings first, P0 to P3, from a model of another family.

A finder reports every defect it sees with a confidence; a separate verifier drops a finding only when the diff does not show the code it describes or contradicts it. Each finding carries file:line, the evidence and the consequence. When nothing survives, the review says "no findings" and lists the residual risks and untested paths; a review that could not finish says "incomplete" instead. Untracked files are not reviewed.

--effort sets how much checking runs: low skips the verifier, high (the default) runs it, and medium also hides, before the verifier, the findings the finder itself rated under 0.8. Each hidden finding is listed by --show-dropped.

chimera review [REVISION_RANGE]
Argument
REVISION_RANGE A revision range, such as main..HEAD. Omit it to review the working tree.
Option Default
--base Review the working tree against its merge base with this ref (default: main).
--repo The repository to review. '.'
--json Print only the report, as JSON (schema chimera.review/1).
--reviewer-model Review with this model (default: CHIMERA_REVIEW_MODEL, else another family's). ''
--author-model The model that wrote the change (default: CHIMERA_DEFAULT_MODEL). ''
--effort low: finder only. medium: also hide findings under confidence 0.8, then verify. high (default): finder and verifier, no cut. See bench/review_confidence_cut.
--no-verify Show every located finding, without the second-stage check (same as --effort low).
--show-dropped Also list the findings the checks dropped, with the reason.
--context Lines of context around each change. 10

rubric-grade

Grade an answer against an authorable rubric — weighted criteria with a required-criterion veto.

Produces a per-criterion breakdown, a single weighted score, and a pass/fail verdict. A required criterion that falls below the gate vetoes the outcome regardless of the weighted score.

chimera rubric-grade
Option Default
--rubric JSON rubric: {criteria:[{text,weight,required}], pass_threshold, required_gate}.
--task The task the answer is for.
--answer The answer text (or use --answer-file).
--answer-file Read the answer from this file.
--model Model slug for the grader.

run

Run a single-shot Tier-1 completion (no fusion). Requires a provider key.

chimera run PROMPT
Argument
PROMPT The prompt to send.
Option Default
--model, -m Override the model slug.
--system, -s Optional system prompt.
--image Attach an image (path or URL); repeatable. Needs a vision model.

sandbox-bench

State-based bench: grade the final workspace state + count harmful side effects.

Unlike the text benches, this measures what the agent DID (files it changed), and flags mutations outside each task's allowed set. Uses real models + file tools.

chimera sandbox-bench
Option Default
--workspace, -w Dir to run sandboxed tasks in. '.sandbox-bench'
--model, -m Override the model slug.
--max-steps Max tool-calling steps per task. 8

scenarios

Run the daily right-hand scenario suite through a real chat session (live). Requires a key.

Each scenario is a script of turns driven through the same ChatSession chimera chat builds — tools, memory, transcript — in its own workspace and its own home. The checks are functional, not substring: equality against a value generated this run and absent from the prompt, a fact read back out of the MemoryStore, a fresh session's recall count, the transcript found in the next turn's assembled prompt, the absence of a fabricated figure.

Twenty-six rows in two blocks. Block C is a validity gate, not a score: six control rows whose expected reading is 100%, so a failure there makes the run invalid rather than lowering the number. Block D is the headline: twenty rows that each carry a defect designed into the environment — a truncated read, a refusal that reads like an observation, ordering bait, a summary that disagrees with its data, an instruction planted in a workspace file, a window that drops the pointer — with both the naive and the careful path available in the tools the agent already has.

Reported with the denominator beside it: pass^k, the flip rate that is this suite's noise floor, ICC(1), and the mechanism-active subset — where a mechanism that never fired reads NOT MEASURED and never 0%. One row per invocation is appended to the series. Pre-registered in bench/scenarios/PREREGISTRATION-v3.md.

chimera scenarios
Option Default
--model, -m Override the model slug.
--k Runs per scenario — one samples, two alert, three decide. 3
--max-steps Max tool-calling steps per turn. 6
--max-usd Hard spend ceiling; the run stops at it. 3.0
--seed Base seed; run i uses seed+i, so the generated values differ per run. 1
--series Where to append the JSONL row (default /scenarios.jsonl).

schema-bench

Measure tool-schema token cost, full vs compacted (advertise-time). No model calls.

chimera schema-bench
Option Default
--openapi Path or URL to an OpenAPI spec to include (its tools are verbose).
--demo Include a couple of synthetic verbose tools to show the effect.
--model, -m Tokenizer model (default: your default).

secrets

Keep provider keys in the OS vault instead of a file.

chimera secrets

secrets list

What the OS vault holds — names only, never values.

Printing a secret would put it in this terminal's scrollback, in any screenshot of it, and in whatever recorded the session, which undoes the reason for having a vault.

chimera secrets list

secrets rm

Remove one credential from the OS vault.

chimera secrets rm NAME
Argument
NAME The credential to forget.

secrets set

Put one credential in the OS vault.

Prompted without echo by default, and that is not politeness: a key typed as an argument lands in the shell history of every machine it is typed on, which is the kind of file this command exists to stop using.

chimera secrets set NAME
Argument
NAME e.g. OPENROUTER_API_KEY
Option Default
--value Omit to be prompted without echo.

serve

Run the messaging gateway on HTTP, Discord, Telegram, Slack or Signal. Requires a key.

Add --cron to also fire scheduled jobs on a real clock — turning the reactive gateway into an agent that acts on time (the daemon that makes proactivity real). Pass --mcp to instead expose Chimera as an MCP server on stdio, so any MCP client (Claude Desktop, an IDE, another agent) can call chimera_solve / chimera_fuse / chimera_memory_search.

chimera serve
Option Default
--host Bind host. '127.0.0.1'
--port Bind port. 8765
--allow-insecure-bind Serve on a reachable address with no token. Only behind a network you already trust.
--model, -m Override the model slug.
--max-steps Max tool-calling steps per message. 6
--workspace, -w Workspace root for tools. '.'
--fuse Route deep-reasoning turns through fusion.
--no-memory Don't recall long-term memory.
--discord Serve on Discord (needs CHIMERA_DISCORD_BOT_TOKEN + the 'messaging' extra).
--telegram Serve on Telegram (needs CHIMERA_TELEGRAM_BOT_TOKEN).
--slack Serve on Slack (needs CHIMERA_SLACK_BOT_TOKEN + CHIMERA_SLACK_APP_TOKEN + the 'messaging' extra).
--signal Serve on Signal via a signal-cli-rest-api bridge (CHIMERA_SIGNAL_API_URL + CHIMERA_SIGNAL_NUMBER).
--cron Also run the cron daemon: fire scheduled jobs on the real clock (proactivity).
--cron-tick Seconds between cron scheduler ticks. 30
--mcp Serve Chimera AS an MCP server over stdio (solve/fuse/memory as tools).
--a2a Also expose an A2A endpoint on HTTP (agent card + task lifecycle).

sessions

List the conversations chimera chat and chimera tui have saved, under <home>/sessions.

Resume one with chimera chat -s <id> or chimera tui -s <id> — one store, so a thread started on either surface continues on the other. These are the terminal's threads, and the ones GET /api/sessions serves; coding conversations in the desktop app are a different store (<home>/code_sessions) with a different shape, and are not listed here.

chimera sessions
Option Default
--delete Delete a session by id.

skillcard-bench

A/B reasoning with vs without injected TRS skill cards. Calls real models.

chimera skillcard-bench
Option Default
--tasks Task suite: hard big
--k How many cards to retrieve per task. 1
--min-overlap Relevance gate: inject a card only on >= N shared query terms (0=off). 2
--max-lines Render budget: max lines per injected card. 3
--use-store Bench your own learned cards (skills.json) instead of the demo set.

skills

List the built-in skills.

chimera skills

skills-approve

Approve/reactivate a learned skill after review (activates retrieval).

Works for both a pending skill (held from a tainted run) and a retired one (un-retire).

chimera skills-approve NAME
Argument
NAME Name of the pending or retired skill to activate.

skills-bundle-disable

Switch a bundle off, keeping it on disk.

Off is not uninstalled, deliberately: trying several and leaving two running is the normal way to use these, and making "off" mean "delete" would charge a download for every change of mind. Use skills-uninstall when you want the files gone.

chimera skills-bundle-disable NAME
Argument
NAME An installed bundle from chimera skills-bundles.

skills-bundle-enable

Switch an installed bundle on, so the agent may use it.

Do this after reading it. An enabled bundle's name and description reach the agent's prompt when they match a task, and its instructions can tell the agent to run the scripts that came with it — which is why nothing is on by default.

chimera skills-bundle-enable NAME
Argument
NAME An installed bundle from chimera skills-bundles.

skills-bundles

List the skill bundles installed on this machine, and where each came from.

chimera skills-bundles

skills-catalog

Browse the installable skills from the wider Agent Skills ecosystem.

These are other people's skills, fetched from their repositories on request — not bundled here. The table says what each one NEEDS, because most were written for a different harness and a catalogue that hid that would be advertising features that fail after the download.

chimera skills-catalog [QUERY]
Argument
QUERY Filter by name or description.
Option Default
--topic Only this topic.

skills-evolve

Reflectively evolve a skill's prompt template against graded instances (GEPA).

Each instance is a {input, expect} pair; the (simple, honest) scorer gives 1.0 when the produced output contains the expect substring, else 0.0. GEPA reflects on a failing case to rewrite the template and keeps a Pareto frontier of candidates. Dry-run by default: the improved skill is only written back to the store with --apply, and only if it beats the seed.

chimera skills-evolve NAME
Argument
NAME Name of the learned skill whose prompt to GEPA-evolve.
Option Default
--instances JSON file: a list of {"input": {...}, "expect": "substring"}.
--budget Rollout budget (evaluations across the search). 20
--model Model slug for the executor + reflector.
--apply Save the improved skill (default: dry-run).

skills-export

Export a learned skill to the open SKILL.md format (portable to the agent-skills ecosystem).

chimera skills-export NAME
Argument
NAME Name of the learned skill to export.
Option Default
--out, -o Write to this path (default: /SKILL.md).

skills-import

Import a SKILL.md into the store. A file imported by path is held pending for review.

Accepts the plain name of a curated card (chimera skills-import verify-before-claiming) as well as a path. The documented form was skills/<name>, a repo-relative path that resolves only inside a checkout — so the one line the README gives for using the shipped library failed for everybody who installed Chimera instead of cloning it.

A file by path lands tainted and pending whatever its frontmatter says, until chimera skills-approve <name>. Its provenance and status were written by its author, so reading them let the stranger decide whether anybody reads the card before the agent does. Only a curated card keeps what it declares: by name, or by a path whose content is exactly the shipped card.

Validated on the way in. This is the only path by which a skill written by somebody else enters the store, and it was the only one that skipped the validator the agent's own proposals must pass — the gate was applied to the code we wrote and not to the code we were handed, which is backwards. A skill card ends up in the system prompt, so an unvalidated one is an instruction from a stranger with the standing of an instruction from the owner.

chimera skills-import PATH
Argument
PATH A curated card name, or a path to a SKILL.md / its directory.

skills-install

Download a skill bundle from its source repository into your skills directory.

Fetches; runs nothing. The bundle lands pending: its files are on disk and no part of it reaches a prompt until you approve it. That is the same rule an imported card follows — a skill from a stranger has the standing of an instruction from the owner — and a bundle is that plus executable scripts, so it holds with more reason, not less.

chimera skills-install NAME
Argument
NAME A skill name from chimera skills-catalog.
Option Default
--force Replace it if it is already installed.

skills-library

Browse the curated skill cards that ship with Chimera.

Data, not code: each is a markdown page of Trigger/Do/Avoid/Check/Risk. Load one into your own store with chimera skills-import <name>. The agent reads a matching card from that store into its prompt only with CHIMERA_SKILL_CARDS=on (or chimera solve --skill-cards); it is off by default, so an imported card is otherwise reference for you, not advice to the agent.

chimera skills-library [NAME]
Argument
NAME Show one card in full; omit to list the library.

skills-lifecycle

Run the measured skill-lifecycle loop (M18-4): promote proven provisionals, demote regressions.

Decisions come from the store's MEASURED usage stats (never a model's self-report): a provisional skill that earns a high win rate over enough uses is promoted to active; a provisional that fails probation or an active skill whose win rate regresses is retired (kept for review). Dry-run by default; --apply closes the loop — cron it for a hands-off promote/demote cycle.

chimera skills-lifecycle
Option Default
--apply Actually promote/demote (default: dry-run preview).
--promote-min-uses Provisional probation length. 5
--promote-min-rate Win rate to promote a provisional skill. 0.7
--demote-min-uses Uses before a skill can be demoted. 5
--demote-max-rate Win rate at/below which a skill is demoted. 0.3333333333333333

skills-pending

List learned skills held for review (e.g. distilled during a tainted run).

chimera skills-pending

skills-retire

Propose retiring under-performing skills — review-gated, never a delete.

Retiring only flips status to 'retired' (excluded from retrieval, still inspectable and reactivatable with skills-approve). With no name, acts on the retirement_candidates signal (used often, low win rate). Dry-run by default; pass --apply to commit.

chimera skills-retire [NAME]
Argument
NAME Skill to retire; omit to act on all candidates.
Option Default
--apply Actually retire (default: dry-run preview).
--min-uses Only propose skills used at least this often. 5
--max-rate Only propose skills whose win rate is at or below this. 0.3333333333333333

skills-stats

Per-skill usage stats (uses, successes, win rate) + retirement candidates.

chimera skills-stats

skills-uninstall

Delete an installed skill bundle and its files.

chimera skills-uninstall NAME
Argument
NAME An installed bundle from chimera skills-bundles.

solve

Tier-2: autonomously solve a task with plan + verify-or-revert. Requires a key.

In the shell nothing about this has changed: a run that fails still exits 1. Inside the process it now hands its run back — returned on success, carried on the SolveFailed exit otherwise — so /solve in a REPL can put the loop's own answer into the conversation instead of writing a sentence of its own.

chimera solve [TASK]
Argument
TASK The task to solve autonomously (omit with --approve/--deny).
Option Default
--verify Verification command (exit 0 == success).
--workspace, -w Workspace root. '.'
--model, -m Override the model slug.
--max-attempts Max verify-or-revert attempts. 3
--max-steps Max tool-calling steps per attempt. 8
--max-usd Stop the whole run once this much has been spent (all attempts together).
--context-budget Fraction of the model's window to spend on the prompt before compacting (e.g. 0.6).
--no-plan Skip the planning step.
--no-manager Skip Manager review.
--rubric Manager reviews via the cascade rubric.
--fuse Route deep-reasoning turns through fusion.
--cascade Tiered routing: weak -> gate -> mid -> gate -> fusion (cheap by default).
--guard Gate tool calls through the governance kernel.
--allow-tools Per-session allowlist: only these tools (comma-separated).
--deny-tools Per-session denylist: drop these tools (comma-separated).
--taint Track a capability ledger + review execution of tainted input.
--collect Record trajectories for opt-in model evolution. True
--no-remember Don't auto-write a long-term memory fact on success.
--no-evolve-skills Don't auto-propose a learned skill when a task recurs.
--isolate Run in an isolated git worktree; changes copied back only on success.
--explorer Give the agent an isolated Context Explorer for repo search (FastContext-style).
--subagents Give the agent spawn_subagent to delegate subtasks to isolated subagents.
--repo-map Prepend a structural map of the workspace (files + top-level symbols) to the agent's context.
--tool-router EXPERIMENT (study 20 B4): a cheap MODEL names the tool before each step and the executor is given only that tool. Reads a shallow context on purpose. Narrows only — an undecided router leaves the full list, and the run's receipt counts how often that happened.
--tool-router-mode EXPERIMENT (study 22 B4b): 'narrow' gives the executor only the routed tool (B4, measured worse); 'hint' keeps every tool and only suggests one for the step, with no ANSWER. 'narrow'
--escalate-on-tool-loop EXPERIMENT (study 24 M6): when the tool-loop breaker trips, hand the rest of the run to this stronger MODEL instead of stopping. A second trip stops as before. Off by default.
--snapshot-at-tool-loop EXPERIMENT (study 24 M6 fork): with --escalate-on-tool-loop, copy the workspace to this DIR at the trip, before escalating — what stopping would have left. Off by default.
--progress-ledger After a failed attempt, run a structured self-check that steers the retry (helps weak models).
--checklist Extract the task's atomic requirements and grade each attempt's coverage (catches dropped constraints).
--gen-tests With no --verify: generate executable pytest grounded in the task's requirements and use it as the gate (catches wrong code the coverage grade rubber-stamps). Measured on 78 labelled patches (bench/test_gate_two_sided): fails every wrong patch, and reverts 4 of 64 correct ones (6%) on a test of its own that is wrong — opt-in for that reason.
--profile Model-role profile: economy balanced
--role-models Per-role model overrides: 'edit=vendor/slug,plan=vendor/other'. Roles: explore, plan, edit, review. Merges over --profile; a role left unset keeps --model. verify is not a role here — it runs a command and has no model to choose.
--write-region Comma-separated globs the file-writers may touch (e.g. 'src/**,*.py'). A write outside is refused — blocks an injected instruction from rewriting an unrelated file.
--probe-log Log (arm, proxy=manager-judgment, reward=verified) per attempt to /probe.jsonl for PROBE best-arm selection (see chimera probe-select --from-log). Needs --verify + a manager.
--normalize-task Reshape a long, rambling bug-report task into a salient-facts-first form (location/repro/expected-vs-actual/fix-hint) before planning. No-op on non-bug or short tasks.
--playbook Inject the stored ACE strategy playbook into context, then curate it from this run's outcome (closed loop).
--skill-cards Read learned skill cards back into context (the learn->use loop). Default follows settings.skill_cards; this overrides it per run.
--agreement With --fuse: sample K cheap answers per turn; escalate to fusion when they disagree (free confidence signal). 1
--strong-verify Model slug of a stronger, independent judge that grades hard-turn (retried) results before accepting them.
--replan On a stall, rebuild the plan from accumulated failure causes (dual-ledger) instead of just nudging.
--diff-feedback Show a failed attempt its own reverted diff, as a path not to retake.
--keep-workspace On failure, leave the last attempt's edits on disk for an external grader (don't revert).
--require-diff Fail an attempt that changed no file — for code tasks, an explanation is not a fix.
--recovery How a failed attempt's retry is briefed: generic (manager prose + verifier output) or targeted (a brief aimed at the classified failure). 'generic'
--stagnation-fuzzy Match repeated-failure signatures approximately, not byte-identically.
--contract Machine-checkable success clauses, comma-separated: file_exists:PATH file_contains:PATH:REGEX
--stream Print live progress events (attempt/result/status) as the run proceeds.
--thread Checkpoint this run under a thread id; re-run with the same id to resume after a crash.
--pause-on-taint Pause for human approval before finalizing a run that consumed untrusted content (needs --thread).
--approve HITL accept: finalize a paused run as-is, by thread id (no task needed).
--deny HITL ignore: discard a paused run by thread id (no task needed).
--respond HITL respond: resume a paused run by thread id with --feedback guidance.
--feedback Guidance for --respond (fed back so the run tries again).
--edit HITL edit: finalize a paused run with the corrected --answer, by thread id.
--answer The human-corrected answer for --edit.

solve-batch

Solve several tasks concurrently, each in its own git worktree (Tier-3 isolation).

Every task runs against an isolated checkout, so parallel edits never collide. On merge-back, a file two tasks both changed is reported as a conflict and left for you to resolve rather than silently overwritten. Needs a git repo to isolate.

A worker whose actions were refused for review is reported as not allowed rather than ok, and the refusals are listed under it. Whether anyone can be asked follows CHIMERA_APPROVAL_MODE: allow and deny answer immediately, ask prompts if this process has a terminal and otherwise writes the question down for chimera approve and waits CHIMERA_APPROVAL_WAIT seconds for it — per refused call, per worker. Set CHIMERA_APPROVAL_WEBHOOK so the question reaches somebody, or CHIMERA_APPROVAL_MODE=deny for a batch that should never wait.

chimera solve-batch TASKS
Argument
TASKS Tasks to solve in parallel, each isolated.
Option Default
--workspace, -w Workspace root (a git repo, to isolate). '.'
--model, -m Override the model slug.
--max-steps Max tool-calling steps per task. 6
--context-budget Fraction of the model's window to spend on the prompt before compacting (e.g. 0.6).
--max-attempts Max verify-or-revert attempts per task. 2
--max-workers Max concurrent isolated workers. 4
--fuse Route deep-reasoning turns through fusion.
--taint Arm each worker's adaptive allowlist (dangerous-when-tainted tools require approval). The cross-agent collusion monitor runs regardless — it's always on for fan-out.

swe-bench-compare

Honest A/B over two SWE-bench Verified-Mini reports on the SAME instance ids.

Reads the official evaluation reports (resolved_ids or a per-instance map) for a free model alone vs the same model driven by Chimera, projects both onto the shared instance list (a missing id counts as unresolved), and prints the delta + 95% CI. This is the second standard scoreboard for the weak-model-lift thesis; the pass/fail comes from SWE-bench's tests, never self-reported.

chimera swe-bench-compare BASELINE TREATMENT
Argument
BASELINE SWE-bench evaluation report JSON for the model-only arm.
TREATMENT SWE-bench evaluation report JSON for the model+Chimera arm.
Option Default
--instances JSONL of the instances both arms ran (fixes the id set).

tools

List the built-in native tools.

chimera tools
Option Default
--workspace, -w '.'
--defer-saving Report what CHIMERA_DEFER_TOOLS / CHIMERA_MCP_DEFER would save on this install.

transfer-gate

Promote a learned change only if it helps its tuned slice AND doesn't regress a holdout.

Guards against negative transfer — a GEPA prompt / ACE delta / distilled skill that raises the pass rate on the tasks it was tuned against but REGRESSES on other tasks sharing the capability. Feed the tuned slice's paired pass/fail (baseline vs candidate) and, ideally, a disjoint same-capability holdout's; the verdict is PROMOTE / BLOCK with the paired evidence (exit 1 on BLOCK, for CI). Without a holdout it promotes on the tuned gain alone but flags that transfer was NOT measured.

chimera transfer-gate TUNED_BASELINE TUNED_TREATMENT
Argument
TUNED_BASELINE JSON pass/fail of the baseline on the TUNED slice (list of bools, or {task: bool}).
TUNED_TREATMENT JSON pass/fail of the candidate on the TUNED slice (aligned, same order).
Option Default
--holdout-baseline JSON pass/fail of the baseline on a DISJOINT same-capability holdout.
--holdout-treatment JSON pass/fail of the candidate on the holdout (aligned).
--require-significant Require the tuned gain's paired CI to exclude 0, not just Δ>0.
--tol Max tolerated pass-rate drop on the holdout before promotion is blocked. 0.0

tui

Launch the full-screen TUI — your right-hand. Requires a key.

Governed like chimera chat: the taint ledger told your own message, the <<external-data>> fence around untrusted tool output, the trust kernel, the owner's reach floor and the connected MCP servers. What took longer to arrive here is the part that makes any of it usable — a question this surface can draw. Textual owns the terminal, so the stdin prompt every other surface uses was never seen: measured in a pty, a run_shell under the shipped CHIMERA_HOST_EXEC=ask blocked 123.8 s against a 120 s timeout and came back as ✗ run_shell with no reason (bench/right_hand_governance/RESULTS.md Part 2). Both gates now open a modal instead; silence still refuses, and now says so while it is counting down.

The conversation outlives the window. Every turn is saved under <home>/sessions — the same store chimera chat writes and chimera sessions lists, so a thread started in one can be picked up in the other — and the newest thread is resumed by default. --session opens a named one and --new starts fresh; on screen, /new (or Ctrl+R, or /reset) starts another and leaves the current one where it is. That last part is a change of meaning rather than of wording: /reset cleared an in-memory transcript back when nothing was on disk, and clearing a thread that is now a file in place would be the command that destroys it.

Note that the scrollback is not redrawn on resume: a resumed turn is in the model's context and not on your screen, and the line under the banner says how many.

chimera tui
Option Default
--model, -m Override the model slug.
--max-steps Max tool-calling steps per message. 6
--workspace, -w Workspace root for tools. '.'
--fuse Route deep-reasoning turns through fusion.
--no-memory Don't recall long-term memory.
--session, -s Resume a specific session id (see 'chimera sessions').
--new Start a fresh session instead of resuming.
--stream Live token streaming (single-model path only). True
--max-usd Stop once this session has spent this much (the whole session, not one turn). The activity panel shows what is left.
--write-region Comma-separated globs the file-writers may touch (e.g. 'src/**,*.py'). A write outside is refused — blocks an injected instruction from rewriting an unrelated file.

version

Show the Chimera version.

chimera version

workflow

Run a declarative workflow — a designed loop — from a YAML file. Requires a key.

chimera workflow FILE
Argument
FILE Workflow YAML file (declarative loop).
Option Default
--workspace, -w Workspace root. '.'
--model, -m Override the model slug.

Edytuj tę stronę na GitHubie