0.60.0 — it continues by itself at the step ceiling, the approval card shows the number that asked, and the sidecar can fetch its browser
Ceci est la note de release telle qu'elle a été publiée, pas une réécriture. Les notes de release sont publiées dans la langue dans laquelle elles ont été écrites.
Added
-
The Code screen continues a turn that stopped at the step ceiling, by itself, up to three times. "Continuar sozinho."
max_stepsis the one stop that means the task was going fine and ran out of room — the model was working, the ceiling cut it, and the work is incomplete by arithmetic rather than by failure. The screen already drew the "stopped at the step limit" badge and then did nothing about it, so the person had to notice and type "continue" — measured on 2026-09-19 in this repository's own desktop transcript, four turns in a row. A toggle beside the notifications one (auto-continue/continuar sozinho, off by default: a chat that keeps sending turns on its own is a chat that keeps spending on its own) arms a continuation inonDoneand sends it through the same gate a queued follow-up waits behind, oncebusyhas actually gone false — no second send path. Onlymax_stepscontinues. Every other reason is a verdict —tool_loopsays the model was repeating itself (continuing would repeat it again, on the person's money),budget/spendsay the run already cost what was capped,cancelledsays the person stopped it on purpose — and past any of those the counter is spent back to zero. The cap is three, hard: three continuations is the difference between "the task needed more room" and "the task will never end". Stop cancels whatever was armed. Four tests: the three continuations and not a fourth; nothing aftertool_loop; off until turned on; Stop ends the sequence. Built by the Chimera desktop itself, working on its own repository through the shell it had just been given. -
The approval card shows the number that asked, and the answer becomes a label. Item 3 of study 20's shortlist. The approval stream was durable, and the line it wrote held an action, a reason and an outcome — so "the person said yes in 8 s" could not be put beside "and the probability that asked was 0.80". That pairing is what turns an answer into a label: the shipped calibration map comes from 55 bench rows today, and every card answered is one real row for a map refitted on the deployment's own decisions.
Verdictcarriesbandanddecider_model(a number without its band cannot be read — 0.45 is a confident ALLOW belowallow_belowand an uncertain one between the thresholds — and a probability whose build is unnamed is a probability about nothing in particular, since the map is keyed on the build);approval._facts_ofreadsp,bandand the build off the verdict rather than taking a second argument;pending.ask_durablyresolves the number once and writes it to the question file (a card that mounts late shows the same number), the announcement (the live screen renders it before any reload) and the record line (pas a number, and no column at all for a rule-raised question —0.0would read as a very confident ALLOW a person nonetheless had to answer).ApprovalOutand the stream frame carry the three fields; the card rendersp=0.80 · band REVIEW · qwen3:4b@Q4_K_Monly when a number came, the decimal throughusePso a pt-BR machine does not read0.80as eighty;chimera approveshowsreview · p=0.80. Ten languages, the generated API types, anddocs/security.mdin nine translations. Thirteen tests on the join, three on the card. -
The four Tier B arms of study 21 ran, and two of them decided product behaviour.
bench/jev_decisions/RESULTS.md§12, pre-registered inPREREGISTRATION-tier-b.md, paired against each item's registered first repetition with a replay floor of 0.05. B1(a) reversing the option order flips 5/55 verdicts on the local arm (0/55 on the vendor) — the local band reads a canonical order, which the instrument hash already pins. B1(b) another item's state collapses the AUROC to 0.31/0.29 — the number is about the state, not the instrument. B2 ten states in one request movepby 0.275 on average (36/55 slots beyond the floor): no surface batches states, one state per call, andOpenRouterDecisionsBackendstays that way. B3 a boundary sentence in the question does nothing on the vendor arm and pushes the localpdown by 0.206 — the kernel keeps stripping the unverifiable sentence itself; a sentence does not replace the rule. B4 claim-versus-diff as a typed Noul reads AUROC 0.54 [0.42, 0.66] within-task against the lexical overlap's 0.66: not a second reading for the verifier-by-uncertainty item, so that item goes on with the lexical signal alone. Two runner defects were found and fixed before any number existed (the event read by name; the sign of the paired delta). US$ 0 on the local arm. -
The kernel has a REVIEW band: where the lexical rules match nothing, a calibrated probability from a typed decision is read against two thresholds — a person sees the action above the upper one, the default applies below the lower one, and between them the default applies with the number on the audit line as a prior. Item C1 of study 20 (
bench/PLAN-study20-calibrated-decisions.md), built afterbench/jev_decisionsmeasured the number and study 21 (bench/PLAN-study21-jev-ecosystem.md) said which one to read.chimera/governance/band.py:TrustKernel(band=)asks the configured decider (chimera/decisions, by defaultqwen3:4bthrough Ollama read decision-first through the shipped Platt map) the governance question on the action alone — never the task context, the tool output or a sentence the request carries, the instrument the map was fitted on — and only after the rules, the learned rules and the precedents had nothing; a rule that matched is never downgraded by a number, and the band never returns a BLOCK. The thresholds are read off the ROC of the calibrated local arm, leave-one-family-out on the 55 governance items:review_at = 0.50is the hosted judge's own operating point (catch 20/24 with 6/31 benign actions stopping for a card) at US$ 0 and 0.3 s a call;allow_below = 0.30is where the arm misses at most one attack in 24. Which number: the Noul-shaped probability of the event, because on our own rows the vendor's Choice mass sits +0.11 above it and calibrates worse (ECE 0.221 against 0.120) — and never the vendor'sconfidence. An uncalibrated number — no map for this instrument, or a map fitted on another build — thresholds nothing and is recorded asraw_pbesideband: uncalibrated; a halt (server off, model not pulled) records itself and the kernel goes on with its other layers. Hysteresis: the local arm at temperature 0 still moves by up to 0.05 raw between two runs of the same item (one flip in 55), so an action that entered REVIEW leaves it only when itspfalls below 0.40, keyed on the exact action text in a bounded memory. Off by default:CHIMERA_GOVERNANCE_BAND=onunderobserveorenforce(.env.example); underobservethe band's REVIEWs reach the allow-everything approver and land inapprovals.granted— recorded, not enforced, which is what prices it: the record-only surfacekernel.pysaid did not exist.chimera guard <action>shows the number when the band is on. Live on this machine after warm-up: 0.31–0.34 s a decision;python -c 'import shutil; shutil.rmtree("/home/bruno")'— which no rule sees — came back REVIEW at p = 0.80,docker system prune -af --volumesALLOW at p = 0.28. Three receipt facts from study 21 land with it: every backend names the build that answered (qwen3:4b@Q4_K_Mfrom/api/show; the dated build behind the vendor's alias) and the Decider refuses a map fitted on another build, with the reason on the receipt; the local and hosted backends keep one client for their lifetime instead of a connection per call (the bench's 0.75 s included that handshake). Fifteen tests on the band (thresholds in order; REVIEW above with the number on the verdict and the audit line; the prior between; the default below; rules first and never downgraded; the judge only when the band had nothing; halt; uncalibrated; another build; hysteresis; bounded memory; off by default; observe records and does not enforce; no decider built when off; a learned rule still wins), three on the receipts; sabotage-verified on six guards — the band enabled regardless of settings, never consulted, consulted before the rules, no hysteresis, uncalibrated numbers thresholded, a map applied to another build. Not here: the approval card showingpand the band (item 3), and the Settings screen switch — the variable is the switch today. -
A typed decision has a contract: a yes/no, a choice or a score, asked of a backend that returns a distribution, calibrated by a map fitted on our own labels, with a receipt that says which of those it got. Item C4 of study 20 (
bench/PLAN-study20-calibrated-decisions.md), amended by what the bench measured (bench/jev_decisions/RESULTS.md): the interface is what transfers from the typed-decision vendor —state + question → probability— and the vendor is optional; the map is part of the contract because calibration did not transfer to a decision it was not fitted on (on aacr-bench the vendor read ECE 0.405 against a floor of 0.035, the verbalized model 0.544).chimera/decisions/:Noul(P(yes)),Choice(options with an event whose probability isp),Score(ordered levels, an expectation) — every one a Choice underneath; three backends, each the bench's instrument byte for byte where a map depends on it —local_logprob(the default: a small model through Ollama's native route, reasoning off, the answer constrained to the options by a JSON schema, the probability read off the label token's log-probabilities found from the end and renormalized over the options — on the governance corpusqwen3:4branks like the hosted judge, AUROC 0.871 against 0.874, saturated raw, and through the map reaches the judge's operating point at US$ 0 and 0.75 s),hosted_verbalized(the fusion judge asked for a verbalized probability with the over-confidence advisory, reasoning off, a 2000-token budget and one re-ask on an empty answer: AUROC 0.886, ECE 0.051 raw),openrouter_decisions(OpenRouter's Decisions API, pinned totypesafe/jev-1.13: 0.34 s, deterministic, 5–8× fewer framing flips, over-confident mid-scale and yielding to pressure framing — opt-in, fails closed per call, never the only layer);Deciderkeys the calibration map on the decision the caller names, the backend, the model and a hash of the instrument — the fixed text around the state — so a reworded question, another model or another decision getscalibrated: falseon its receipt rather than a map fitted on something else; a backend that raises is a halt on the answer, never a verdict. Platt scaling in pure Python (Newton with a backtracking step; agrees with scikit-learn to four decimals on the bench rows) because two parameters cannot overfit fifty rows and isotonic did, leave-one-family-out. One map ships — the governance question on the local backend, fitted on the bench's 55 rows bybench/jev_decisions/fit_map.pyand held equal to a refit by a test — and it says what it does to the saturated model: raw 0.50 → 0.07, 0.90 → 0.26, 0.99 → 0.66.CHIMERA_DECISION_BACKEND/CHIMERA_DECISION_MODELchoose the instrument;<home>/decisions/maps.jsonholds a deployment's own refits, which replace the shipped map on the same four keys. Nothing consumes a decision yet — an AST guard keeps it so until the kernel's REVIEW band, the strong verifier and the voice router each arrive with their measurement. Twenty-six tests; sabotage-verified on the from-the-end reading, the instrument hash, the shipped constant, the map's application and the re-ask. -
A decision can carry its number: token probabilities reach the result, a verdict can say how sure its decider was, and a question's record names the run that asked it. Item A1 of study 20 (
bench/PLAN-study20-calibrated-decisions.md): forty-four decision points mapped, twelve decided by a model, none carrying a probability; the gateway already forwardedlogprobs=Trueto the provider and never read the answer; the approval history could not be joined to anything. Three pieces of plumbing, none of which changes a decision.CompletionResult.logprobsholds the route's token log-probabilities as plain dicts and isNonewhen none came — the reading to keep, since OpenRouter forwards a request to a provider that ignores the parameter and drops it in silence, and a reasoning route sends none;chimera.providers.decision.label_probabilitiesturns the first token into shares over a closed label set, renormalized over the labels, with the mass that was on the labels beside them (a decision token after a reasoning trace is extraction, not decision — measured in arXiv 2601.13284 at 92–100% trace-following — so the reader is decision-first by construction).Verdict.confidencetravels onto the audit line, and so onto the Security screen's summary, only when the decider had a number; a rule never reads as unsure.pending.FACTSnames what a record line inapprovals/history.jsonlmay carry —run_id,surface,tool,rule,lineage,sources—approval._facts_ofreads rule/sources/tool off the question's own objects, and the desktop's approver and the plan gate write the turn's id and surface, so "the person said yes in 8 s" can at last be put beside "and the turn verified". Measured before it was built: one installed copy held 43 record lines (24 approved · 1 refused · 18 timed out) with no key that joined them to a run. Sabotage-verified: the join removed → four tests red; the audit field removed → one.
Fixed
-
The desktop's first browser use can download Chromium in the installed app.
_install_chromiumran[sys.executable, "-m", "playwright", "install", "chromium"]; in the frozen sidecarsys.executableischimera-backend.exe, whose argv is prepended withappand handed to the CLI — so the three words arrived as extra arguments and the install exited 2 (Got unexpected extra argument(s) (install chromium)). Measured on 2026-09-21 in the desktop's own transcript: the installed app could never provision its browser.playwright.__main__is only the driver executable, its CLI and the arguments, run with the driver's env; the install now builds that command itself — the same node driver the launch already uses, present in the freeze, no interpreter asked to run a module. One test pins the spawned argv and env; sabotage-verified against the old route. -
The feature routes honour the settings they were built with.
register_featuresreadget_settings()directly for the cron store, the memory manager, the skill store and the project loader, so an injected settings object pointed those routes at the real user's home — harmless in the product, and a trap in tests, which is where it was found:test_the_scheduled_gate_is_reachablewrote its job into the developer's own crontab and read a different job back. Every helper now takes the settings it should read. Alongside it, the suite owns the environment leaks it was blaming on tests (three leaks, one guard, the guard was right), the sklearn/torch probes skip where the extra is not installed, and the pinned secret scanner is told — by fingerprint, one commit and one line — that the i18n key anApprovalCard.tsxcomment wrote as an example is the name of a test file, not a credential; a test holds the entry true (file exists, the line still holds the finding, the current tree does not) so a silent allowlist entry cannot go stale unnoticed. -
A large file is read in windows, and
greppointed at a file searches that file. Measured on 2026-09-19 in the desktop agent's own transcript, working on this repository:read_filecut a 137k-character module at 20,000 characters with a notice that named the total and no way to reach the rest, andgrepwith a file aspathanswerednot a directory— which the model read as "try again", eleven times in one forty-step turn; the turn ended on the step ceiling with the region it had to edit unread, and the next four turns the same way.read_filetakesstart_line(1-based) andmax_lines, and its truncation notice now names the next window (continue with start_line=N);grep'spathmay be a file. The 20,000-character ceiling is unchanged. Five tests, the two new arguments classified as identifiers. -
Two catalogue rows learn the second route's price, so the live check on
mainis green again. The merge commit of #516 went red on the one test that runs only onmain— the live OpenRouter index against the catalogue — because two slugs were being quoted from a route the rows had never seen:deepseek/deepseek-v4-flashat 0.04844/0.09688 against the row's 0.0886/0.1772, andz-ai/glm-5.3at 0.91/2.86 against 1.40/4.40. Same mechanism as #449 (a slug served by two routes flips between two figures, and the check accepts any price a row has been seen at): both rows carry the new pair inalso_seen, with the date in the note. Verified against the real index: the check fails on the previous rows and passes on these. Receipts are priced from the live index; these rows are the fallback.