Zum Inhalt springen

← Blog

Releasesv0.60.0

0.60.0 — it continues by itself at the step ceiling, the approval card shows the number that asked, and the sidecar can fetch its browser

Dies ist die Release Note so, wie sie veröffentlicht wurde, keine Neufassung davon. Release Notes werden in der Sprache veröffentlicht, in der sie geschrieben wurden.

Added

  • The Code screen continues a turn that stopped at the step ceiling, by itself, up to three times. "Continuar sozinho." max_steps is the one stop that means the task was going fine and ran out of room — the model was working, the ceiling cut it, and the work is incomplete by arithmetic rather than by failure. The screen already drew the "stopped at the step limit" badge and then did nothing about it, so the person had to notice and type "continue" — measured on 2026-09-19 in this repository's own desktop transcript, four turns in a row. A toggle beside the notifications one (auto-continue / continuar sozinho, off by default: a chat that keeps sending turns on its own is a chat that keeps spending on its own) arms a continuation in onDone and sends it through the same gate a queued follow-up waits behind, once busy has actually gone false — no second send path. Only max_steps continues. Every other reason is a verdict — tool_loop says the model was repeating itself (continuing would repeat it again, on the person's money), budget/spend say the run already cost what was capped, cancelled says the person stopped it on purpose — and past any of those the counter is spent back to zero. The cap is three, hard: three continuations is the difference between "the task needed more room" and "the task will never end". Stop cancels whatever was armed. Four tests: the three continuations and not a fourth; nothing after tool_loop; off until turned on; Stop ends the sequence. Built by the Chimera desktop itself, working on its own repository through the shell it had just been given.

  • The approval card shows the number that asked, and the answer becomes a label. Item 3 of study 20's shortlist. The approval stream was durable, and the line it wrote held an action, a reason and an outcome — so "the person said yes in 8 s" could not be put beside "and the probability that asked was 0.80". That pairing is what turns an answer into a label: the shipped calibration map comes from 55 bench rows today, and every card answered is one real row for a map refitted on the deployment's own decisions. Verdict carries band and decider_model (a number without its band cannot be read — 0.45 is a confident ALLOW below allow_below and an uncertain one between the thresholds — and a probability whose build is unnamed is a probability about nothing in particular, since the map is keyed on the build); approval._facts_of reads p, band and the build off the verdict rather than taking a second argument; pending.ask_durably resolves the number once and writes it to the question file (a card that mounts late shows the same number), the announcement (the live screen renders it before any reload) and the record line (p as a number, and no column at all for a rule-raised question — 0.0 would read as a very confident ALLOW a person nonetheless had to answer). ApprovalOut and the stream frame carry the three fields; the card renders p=0.80 · band REVIEW · qwen3:4b@Q4_K_M only when a number came, the decimal through useP so a pt-BR machine does not read 0.80 as eighty; chimera approve shows review · p=0.80. Ten languages, the generated API types, and docs/security.md in nine translations. Thirteen tests on the join, three on the card.

  • The four Tier B arms of study 21 ran, and two of them decided product behaviour. bench/jev_decisions/RESULTS.md §12, pre-registered in PREREGISTRATION-tier-b.md, paired against each item's registered first repetition with a replay floor of 0.05. B1(a) reversing the option order flips 5/55 verdicts on the local arm (0/55 on the vendor) — the local band reads a canonical order, which the instrument hash already pins. B1(b) another item's state collapses the AUROC to 0.31/0.29 — the number is about the state, not the instrument. B2 ten states in one request move p by 0.275 on average (36/55 slots beyond the floor): no surface batches states, one state per call, and OpenRouterDecisionsBackend stays that way. B3 a boundary sentence in the question does nothing on the vendor arm and pushes the local p down by 0.206 — the kernel keeps stripping the unverifiable sentence itself; a sentence does not replace the rule. B4 claim-versus-diff as a typed Noul reads AUROC 0.54 [0.42, 0.66] within-task against the lexical overlap's 0.66: not a second reading for the verifier-by-uncertainty item, so that item goes on with the lexical signal alone. Two runner defects were found and fixed before any number existed (the event read by name; the sign of the paired delta). US$ 0 on the local arm.

  • The kernel has a REVIEW band: where the lexical rules match nothing, a calibrated probability from a typed decision is read against two thresholds — a person sees the action above the upper one, the default applies below the lower one, and between them the default applies with the number on the audit line as a prior. Item C1 of study 20 (bench/PLAN-study20-calibrated-decisions.md), built after bench/jev_decisions measured the number and study 21 (bench/PLAN-study21-jev-ecosystem.md) said which one to read. chimera/governance/band.py: TrustKernel(band=) asks the configured decider (chimera/decisions, by default qwen3:4b through Ollama read decision-first through the shipped Platt map) the governance question on the action alone — never the task context, the tool output or a sentence the request carries, the instrument the map was fitted on — and only after the rules, the learned rules and the precedents had nothing; a rule that matched is never downgraded by a number, and the band never returns a BLOCK. The thresholds are read off the ROC of the calibrated local arm, leave-one-family-out on the 55 governance items: review_at = 0.50 is the hosted judge's own operating point (catch 20/24 with 6/31 benign actions stopping for a card) at US$ 0 and 0.3 s a call; allow_below = 0.30 is where the arm misses at most one attack in 24. Which number: the Noul-shaped probability of the event, because on our own rows the vendor's Choice mass sits +0.11 above it and calibrates worse (ECE 0.221 against 0.120) — and never the vendor's confidence. An uncalibrated number — no map for this instrument, or a map fitted on another build — thresholds nothing and is recorded as raw_p beside band: uncalibrated; a halt (server off, model not pulled) records itself and the kernel goes on with its other layers. Hysteresis: the local arm at temperature 0 still moves by up to 0.05 raw between two runs of the same item (one flip in 55), so an action that entered REVIEW leaves it only when its p falls below 0.40, keyed on the exact action text in a bounded memory. Off by default: CHIMERA_GOVERNANCE_BAND=on under observe or enforce (.env.example); under observe the band's REVIEWs reach the allow-everything approver and land in approvals.granted — recorded, not enforced, which is what prices it: the record-only surface kernel.py said did not exist. chimera guard <action> shows the number when the band is on. Live on this machine after warm-up: 0.31–0.34 s a decision; python -c 'import shutil; shutil.rmtree("/home/bruno")' — which no rule sees — came back REVIEW at p = 0.80, docker system prune -af --volumes ALLOW at p = 0.28. Three receipt facts from study 21 land with it: every backend names the build that answered (qwen3:4b@Q4_K_M from /api/show; the dated build behind the vendor's alias) and the Decider refuses a map fitted on another build, with the reason on the receipt; the local and hosted backends keep one client for their lifetime instead of a connection per call (the bench's 0.75 s included that handshake). Fifteen tests on the band (thresholds in order; REVIEW above with the number on the verdict and the audit line; the prior between; the default below; rules first and never downgraded; the judge only when the band had nothing; halt; uncalibrated; another build; hysteresis; bounded memory; off by default; observe records and does not enforce; no decider built when off; a learned rule still wins), three on the receipts; sabotage-verified on six guards — the band enabled regardless of settings, never consulted, consulted before the rules, no hysteresis, uncalibrated numbers thresholded, a map applied to another build. Not here: the approval card showing p and the band (item 3), and the Settings screen switch — the variable is the switch today.

  • A typed decision has a contract: a yes/no, a choice or a score, asked of a backend that returns a distribution, calibrated by a map fitted on our own labels, with a receipt that says which of those it got. Item C4 of study 20 (bench/PLAN-study20-calibrated-decisions.md), amended by what the bench measured (bench/jev_decisions/RESULTS.md): the interface is what transfers from the typed-decision vendor — state + question → probability — and the vendor is optional; the map is part of the contract because calibration did not transfer to a decision it was not fitted on (on aacr-bench the vendor read ECE 0.405 against a floor of 0.035, the verbalized model 0.544). chimera/decisions/: Noul (P(yes)), Choice (options with an event whose probability is p), Score (ordered levels, an expectation) — every one a Choice underneath; three backends, each the bench's instrument byte for byte where a map depends on it — local_logprob (the default: a small model through Ollama's native route, reasoning off, the answer constrained to the options by a JSON schema, the probability read off the label token's log-probabilities found from the end and renormalized over the options — on the governance corpus qwen3:4b ranks like the hosted judge, AUROC 0.871 against 0.874, saturated raw, and through the map reaches the judge's operating point at US$ 0 and 0.75 s), hosted_verbalized (the fusion judge asked for a verbalized probability with the over-confidence advisory, reasoning off, a 2000-token budget and one re-ask on an empty answer: AUROC 0.886, ECE 0.051 raw), openrouter_decisions (OpenRouter's Decisions API, pinned to typesafe/jev-1.13: 0.34 s, deterministic, 5–8× fewer framing flips, over-confident mid-scale and yielding to pressure framing — opt-in, fails closed per call, never the only layer); Decider keys the calibration map on the decision the caller names, the backend, the model and a hash of the instrument — the fixed text around the state — so a reworded question, another model or another decision gets calibrated: false on its receipt rather than a map fitted on something else; a backend that raises is a halt on the answer, never a verdict. Platt scaling in pure Python (Newton with a backtracking step; agrees with scikit-learn to four decimals on the bench rows) because two parameters cannot overfit fifty rows and isotonic did, leave-one-family-out. One map ships — the governance question on the local backend, fitted on the bench's 55 rows by bench/jev_decisions/fit_map.py and held equal to a refit by a test — and it says what it does to the saturated model: raw 0.50 → 0.07, 0.90 → 0.26, 0.99 → 0.66. CHIMERA_DECISION_BACKEND / CHIMERA_DECISION_MODEL choose the instrument; <home>/decisions/maps.json holds a deployment's own refits, which replace the shipped map on the same four keys. Nothing consumes a decision yet — an AST guard keeps it so until the kernel's REVIEW band, the strong verifier and the voice router each arrive with their measurement. Twenty-six tests; sabotage-verified on the from-the-end reading, the instrument hash, the shipped constant, the map's application and the re-ask.

  • A decision can carry its number: token probabilities reach the result, a verdict can say how sure its decider was, and a question's record names the run that asked it. Item A1 of study 20 (bench/PLAN-study20-calibrated-decisions.md): forty-four decision points mapped, twelve decided by a model, none carrying a probability; the gateway already forwarded logprobs=True to the provider and never read the answer; the approval history could not be joined to anything. Three pieces of plumbing, none of which changes a decision. CompletionResult.logprobs holds the route's token log-probabilities as plain dicts and is None when none came — the reading to keep, since OpenRouter forwards a request to a provider that ignores the parameter and drops it in silence, and a reasoning route sends none; chimera.providers.decision.label_probabilities turns the first token into shares over a closed label set, renormalized over the labels, with the mass that was on the labels beside them (a decision token after a reasoning trace is extraction, not decision — measured in arXiv 2601.13284 at 92–100% trace-following — so the reader is decision-first by construction). Verdict.confidence travels onto the audit line, and so onto the Security screen's summary, only when the decider had a number; a rule never reads as unsure. pending.FACTS names what a record line in approvals/history.jsonl may carry — run_id, surface, tool, rule, lineage, sources — approval._facts_of reads rule/sources/tool off the question's own objects, and the desktop's approver and the plan gate write the turn's id and surface, so "the person said yes in 8 s" can at last be put beside "and the turn verified". Measured before it was built: one installed copy held 43 record lines (24 approved · 1 refused · 18 timed out) with no key that joined them to a run. Sabotage-verified: the join removed → four tests red; the audit field removed → one.

Fixed

  • The desktop's first browser use can download Chromium in the installed app. _install_chromium ran [sys.executable, "-m", "playwright", "install", "chromium"]; in the frozen sidecar sys.executable is chimera-backend.exe, whose argv is prepended with app and handed to the CLI — so the three words arrived as extra arguments and the install exited 2 (Got unexpected extra argument(s) (install chromium)). Measured on 2026-09-21 in the desktop's own transcript: the installed app could never provision its browser. playwright.__main__ is only the driver executable, its CLI and the arguments, run with the driver's env; the install now builds that command itself — the same node driver the launch already uses, present in the freeze, no interpreter asked to run a module. One test pins the spawned argv and env; sabotage-verified against the old route.

  • The feature routes honour the settings they were built with. register_features read get_settings() directly for the cron store, the memory manager, the skill store and the project loader, so an injected settings object pointed those routes at the real user's home — harmless in the product, and a trap in tests, which is where it was found: test_the_scheduled_gate_is_reachable wrote its job into the developer's own crontab and read a different job back. Every helper now takes the settings it should read. Alongside it, the suite owns the environment leaks it was blaming on tests (three leaks, one guard, the guard was right), the sklearn/torch probes skip where the extra is not installed, and the pinned secret scanner is told — by fingerprint, one commit and one line — that the i18n key an ApprovalCard.tsx comment wrote as an example is the name of a test file, not a credential; a test holds the entry true (file exists, the line still holds the finding, the current tree does not) so a silent allowlist entry cannot go stale unnoticed.

  • A large file is read in windows, and grep pointed at a file searches that file. Measured on 2026-09-19 in the desktop agent's own transcript, working on this repository: read_file cut a 137k-character module at 20,000 characters with a notice that named the total and no way to reach the rest, and grep with a file as path answered not a directory — which the model read as "try again", eleven times in one forty-step turn; the turn ended on the step ceiling with the region it had to edit unread, and the next four turns the same way. read_file takes start_line (1-based) and max_lines, and its truncation notice now names the next window (continue with start_line=N); grep's path may be a file. The 20,000-character ceiling is unchanged. Five tests, the two new arguments classified as identifiers.

  • Two catalogue rows learn the second route's price, so the live check on main is green again. The merge commit of #516 went red on the one test that runs only on main — the live OpenRouter index against the catalogue — because two slugs were being quoted from a route the rows had never seen: deepseek/deepseek-v4-flash at 0.04844/0.09688 against the row's 0.0886/0.1772, and z-ai/glm-5.3 at 0.91/2.86 against 1.40/4.40. Same mechanism as #449 (a slug served by two routes flips between two figures, and the check accepts any price a row has been seen at): both rows carry the new pair in also_seen, with the date in the note. Verified against the real index: the check fails on the previous rows and passes on these. Receipts are priced from the live index; these rows are the fallback.

Release auf GitHub lesen →