v0.64.0
Это описание выпуска как оно опубликовано, а не его пересказ. Описания выпусков публикуются на том языке, на котором были написаны.
Added
- The dynamic screen, phase 7: any panel of the right panel opens in a window of its own.
- How: "To a window of its own" in a panel's move menu. The panel leaves its dock while the window is open and comes back however the window goes: its own "Bring back" button, the tray in the status bar (which lists it as "in a window"), closing it, or the window vanishing without a word (checked every second). Closing the main window closes them all, and the app quits as before.
- The window is the same page, asked by
/?float=<panel>to draw one panel. The agent's state reaches it over aBroadcastChannel; the panels that ask the server themselves need nothing. - Which panels float is not stored. A relaunch finds every panel in its dock; undo, the default layout and the server copy never meet a window.
- Native side: this runtime refused every new window until now. The main window now has one
handler, and it opens a window for exactly one address: this backend's origin, the root path and a
single
floatparameter naming a panel. Everything else is still refused. The new windows get no IPC: the capability file namesmainalone, and a test fails if it ever grants more.
- The dynamic screen, phase 6: the layout is kept by the server too, and the editor's sidebar joins in.
- Why: the layout lived only in the webview's storage, which a reinstall or a cleared WebView2 profile wipes. The owner chose to keep it where the project list already lives.
- Server:
GETandPUT /api/ui/layoutkeep it as one JSON object underCHIMERA_HOME(chimera/core/ui_layout.py): an object only, at most 64 KB, replaced atomically. The server does not check its shape; the client owns the model and reads anything it does not recognise as its default. - Client (
lib/layout/sync.tsx, mounted inmain.tsx):- the first time, the local layout goes up;
- after that the server's is applied at start (not as a step to undo) and each change is sent once after a short pause;
- a change made before the server answered wins;
- a stored value of another version is no answer, so the local layout is kept and sent;
- with no server, the local copy keeps working.
- The editor's sidebar follows the left region like the conversation list: it hides from its own button, leaves a tab on its edge and slides back in.
- Opening the desktop app again brings the open window forward. A second launch hands over to the running app instead of starting a second backend on the same data folder.
- The dynamic screen, phase 5: maximise any panel, focus mode, and layouts one command away.
- Maximise: the file viewer and the tools, fusion, background-jobs and machine panels maximise from a button in their header, or with ⌘⇧M / Ctrl+Shift+M on the panel that holds focus. Escape restores from anywhere, unless an open menu or dialog already took it. A maximised viewer takes the whole row; closing the file restores the layout, so the next file does not open maximised.
- Focus mode: a button in the status bar (and ⌘⇧F / Ctrl+Shift+F, and the palette) puts the rail and both side regions away and shrinks the composer's settings to their line; the same button puts back exactly what was there.
- Layouts: "review" (the file viewer as wide as it goes, the list away) and "monitor" (the right panel as wide as it goes) in the command palette, plus "Save this layout as mine" and "Apply my layout".
- ⌘⇧A / Ctrl+Shift+A goes to the approval waiting in the conversation, minimised or not.
- The shifted chords are read by the physical key and leave the unshifted ones alone: Ctrl+A still selects, Ctrl+F still finds. Nine translation keys in all ten languages.
- The dynamic screen, phase 4: the right panel's sections move between the right panel, the left sidebar and a
bottom dock.
- What moves: tools, tokens, memory, fusion, background jobs and this machine are each a panel
(
shell/Dock.tsx,shell/panels.tsx) with a drag handle, a "Move to" menu, minimise and close (with an Undo, and listed in the hidden-panels tray). A bottom dock appears when a panel is moved there, and only then; while a panel is being dragged, an empty zone shows where it can land. - Dragging by keyboard: Space or Enter on the handle picks a panel up, the arrow keys move it (between
zones too), Space drops it and Escape cancels; each step is announced to a screen reader in the reader's
language. A pointer drag needs a few pixels of movement first, so a click on the handle stays a click.
This is
@dnd-kit(core, sortable, utilities), the dependency the owner approved for this phase. - What does not move: the agent's state line at the top of the right panel is no longer in the list of panels, because it is the agent's state, which never hides; the conversation list stays home.
- The composer's settings (who runs, cost profile, model) minimise to one line of chips. The posture note, with the no-sandbox warning, is outside them and stays on screen.
- A panel whose content has nothing to say (Fusion without a fused turn, no background jobs) hides its frame too. The machine panel keeps its old title. Eleven translation keys in all ten languages.
- What moves: tools, tokens, memory, fusion, background jobs and this machine are each a panel
(
- Turns running at once warn about what they spend together. Each turn already warned at its own US$ 1; five running at once could spend five times that without a word. When two or more running turns together cross a multiple of the warning amount, each says so. It is a warning, not a stop; a ceiling the person typed is still that turn's own.
- A conversation can open in a window of its own, so two are worked at once. "Open in a new window" in the conversation's header draws that one conversation in its own window, in its own project, with its own composer. The window reads the layout and writes none of it, offers no further window, and a second click focuses it. Checked live in the desktop shell.
- The dynamic screen, phase 3: every card of the conversation minimises, closes and comes back.
- What: the tool list, the task list, warnings, the browser view, each file's changes, the verification
verdict, the receipt, a failed turn's error and the approval card each carry three controls in their corner
(
code/CardChrome.tsx): minimise to one line, "always minimise this kind" (kept in the layout), and close. - Closing is for this screen only: the card disappears from that turn with an Undo in a toast, and the turn shows "Hidden in this turn: N · show" where it was. Nothing is deleted or stored; reopening the conversation shows its cards.
- Three kinds minimise and never close: the approval card, spend and limit warnings, and a failed turn's error. Their close button stays, disabled, with the reason as its tooltip. A new approval always opens, whatever the last one was left as.
- The toast gained one optional action, for undoing what the person just did, and
useOptionalToast, which does nothing where no provider is mounted: the Undo is a shortcut on top of the chip, not the only way back. Twenty-two translation keys in all ten languages.
- What: the tool list, the task list, warnings, the browser view, each file's changes, the verification
verdict, the receipt, a failed turn's error and the approval card each carry three controls in their corner
(
- The dynamic screen, phase 2: drag the widths of the conversation list, the right panel and the file viewer.
- How: a thin line between the columns (
shell/Splitter.tsx) that can be dragged, moved 16 px at a time with the arrow keys once focused, and put back to its starting width with Home or a double click. It is the WAI-ARIA window-splitter pattern: a focusableseparatorthat states the width it controls and its range. - The widths are the layout's, clamped there (the list 180–420 px, the right panel 220–480, the viewer 280–900) and kept across launches. One drag is one step to undo, and a second drag is a second step.
- Side by side only: below 1024 px the Code screen stacks, and there is no width to drag between rows.
- Not
react-resizable-panels. The plan named it and the owner approved it; building the phase showed it does not fit, because it sizes sibling panels inside one group while the right panel lives in the shell and the list inside the Code screen, and the layout already keeps the widths, their limits and their storage. What it would have bought is the one small file above.@dnd-kitfor phase 4 is unaffected. - Three translation keys in all ten languages.
columns-can-shrink.test.tsnow reads the widths from the layout, with the reason in the test: the roles did not change, only where the width comes from.
- How: a thin line between the columns (
- The dynamic screen, phase 1: hide and show the left sidebar, the right panel and the screen rail.
- How: a button in the conversation list's header and in the right panel's header, ⌘B / Ctrl+B for the left and ⌘⌥B / Ctrl+Alt+B for the right (read by the physical key, so AltGr on a Brazilian keyboard still works), and a command per region in the command palette, worded for what it will do now.
- The way back is where the thing went: a hidden region leaves a narrow tab on its edge, and hiding hands focus to that tab so a keyboard user is not dropped on the page. The hidden-panels tray lists it too.
- The agent's state does not go with the right panel: the status bar keeps the state and Stop.
- Motion: a region that comes back slides in from its own edge (
duration-3, transform and opacity only), and only when it comes back, not every time a screen opens. Reduced motion collapses it to 1 ms. - The choice is kept across launches, in the layout from phase 0. Six translation keys in all ten languages.
- Autonomous runs in parallel, one per project. A run in one project no longer blocks starting a run, or the verdict's fix, in another. A second run in the same project is still refused, and the status bar names the latest run with a count of the others.
- The dynamic screen, phase 0: one layout model, a way back to anything hidden, and "Restore default layout".
- What it is: the first of eight phases of the plan the owner approved on 2026-09-29, a screen where anything can be minimised, maximised, closed, dragged, resized and brought back. This phase is the model the others draw from, and changes nothing on screen by itself.
- The model: the Code screen's layout is one serialisable value (
lib/layout/model.ts), changed only by a pureapplyLayout(layout, action). That gives undo, "restore default" and tests that need no screen. Closing keeps a panel's place, so reopening puts it back where it was. Focus mode remembers the layout it replaced and returns to it exactly. - What never disappears: the approval card, spend and limit warnings and a failed turn's error only
minimise; Stop and the status bar are not panels, so no action reaches them. The rules live in
applyLayout, not in the buttons, and a refused action returns the same object. - Stored locally under
chimera.layout.v1.code. A stored layout can never break the screen: anything not recognised becomes the default, a panel a newer build adds takes its default place, and a storage that throws costs the layout and nothing else. - Visible now: a hidden-panels tray in the status bar that lists what is hidden with "Show", "Show all"
and "Restore default layout". It renders nothing while nothing is hidden, following the rule
PendingApprovalsalready set, so "Restore default layout" is also in the command palette. - Twenty new translation keys, in all ten languages.
DESIGN.mdgains a "Dynamic layout" section.
- A coding turn that is still running can be found again: in the sidebar, and in a conversation reopened mid-turn.
- What was wrong (measured live, 2026-09-29): a turn keeps running on the server when the screen
that started it goes away, and its frames are kept. But the conversation is stored when the agent
finishes, so for as long as the turn ran its session read back empty (
exchanges: []) and was missing from the list. Someone who moved to another conversation and came back saw an idle one, with nothing to say that a task was working in it. - Now: the server keeps an in-memory registry of the turns running.
GET /api/code/turns/runninglists them; each row of the session list carriesrunning(and a conversation whose first turn is still running is listed although its file does not exist yet); the session read carriesrunning_turn. The sidebar marks the conversations that are working. A conversation reopened mid-turn follows the turn through its live stream from its opening frame, shows it as working (a message typed meanwhile waits behind it, as it does behind a turn the screen started), and comes back when it ends. - The subtle part: the agent's run saves the transcript and the turn goes on verifying after that, so for a while the stored file holds the exchange of a turn that is still running. Each turn counts its own transcript writes, and a session read is trusted only when the count did not move under it; the screen then drops the stored copy and replays the turn, so the exchange is drawn once.
- What did not change: stopping a followed turn stops watching it, exactly what it means for a turn the screen started; the turn goes on running on the server. Nothing here can stop a turn from another screen.
- The registry is in memory on purpose: a turn is a thread of this process, so a restart ends every
turn and a record on disk would go on saying "running". One new translation key,
code.sessions.running, in all ten languages.
- What was wrong (measured live, 2026-09-29): a turn keeps running on the server when the screen
that started it goes away, and its frames are kept. But the conversation is stored when the agent
finishes, so for as long as the turn ran its session read back empty (
- A credential in a decision's state no longer reaches a hosted backend or the decision log.
- What was wrong: nothing in
chimera/decisions/used the redaction net. The decision log wrote the first 500 characters of the state raw, and both hosted backends sent the state whole. The REVIEW band's state is a shell command, which is exactly where a bearer token sits. - Now: the state goes through
chimera.core.redact(the net that already keeps secrets out of the trace) before a hosted backend reads it, and before the log writes it. The log redacts first and cuts second, so a token straddling the 500-character cut is masked rather than left as a fragment. The state hash is still of the original, so a refit still joins rows on it. - What did not change: the local backend, which sends nothing off the machine.
- A new page,
docs/decisions-redaction.md, says per surface what the state carries, what each backend sees, and what is stored. It also records what is not done, on purpose: opaque substitution of paths (a state without its paths is a state the danger question cannot judge) and a "coarse features only" mode. - Study 27, phase 3, third slice.
- What was wrong: nothing in
- Drift alarms on the Decisions screen and in
chimera decisions report, computed from the log alone.- What they say: the serving build behind one instrument changed (
model_changed, and the calibration map is keyed on the build); the last 50 raw answers moved away from the rest (answer_drift: a population-stability index of at least 0.2 and a chi-square test under 0.01, both required); and 25% or more of the recent calibrated answers sit within 0.03 of a cut of the band (near_threshold, a region that flaps). - What they do not do: gate anything. They annotate. A refit that follows one can now be justified by the log instead of a hunch.
- How:
chimera/decisions/drift.py, pure Python (no scipy in this project; the chi-square survival function is the incomplete gamma, checked against published table values). The screen shows the panel only when there is something to say, in ten languages. - Study 27, phase 3, first slice. The rest of the phase (the spend and rate gate, the decisions read model, the redaction page) follows as separate changes.
- What they say: the serving build behind one instrument changed (
- A run can now warn without stopping. The loop's limits were silent until they were a stop; there is a third thing now, a
notice.- What it says:
tool_loop_warnwhen the same call has repeated three times (the breaker used to compute this level and drop it),steps_lowtwo steps before the step limit, andcompactedwhen the conversation was compacted to keep going. - Where it shows: the desktop turn stream gets a
noticeframe, drawn as one warning line per code under the task list, in all ten languages. The terminal REPL prints the same warnings live. - What it does not do: change how any run ends. This is the channel only: the same run ends the same way with or without a listener, and a listener that raises is ignored. An agent written before the channel is never handed the keyword.
- What it says:
- Two new question-lint rules.
duplicate_criteria(error) refuses a question whose options share one criterion, since nothing in the rubric tells them apart.negated_true(warning) flags a yes/no question whose "true" criterion is itself a negation, the double negative that breaks P(X) + P(not X) = 1. - A gate in front of every hosted decision ask: it waits while a budget is busy and refuses, naming the gate, when waiting will not help.
- What it does: a sliding 60-second window of requests and tokens. At 80% of a configured
CHIMERA_DECISION_RPMorCHIMERA_DECISION_TPMthe next ask waits for the oldest entry to age out, up to 30 seconds, and past that it is refused instead of hanging a decision. A429with aRetry-Afterholds the gate for that long.CHIMERA_DECISION_DAILY_USDis a ceiling over what the decision log says today's hosted answers cost. - The refusal is designed: a refused ask is a halt whose receipt (and log line) carries
gate: "budget"or"rate", never an answer and never silence. What the REVIEW band does with that is a separate change. - An unpriced answer is unknown, not free. The hosted backend now reports
usdas unknown when the model has no price, where it used to report 0.0. With a daily ceiling set, an unknown makes the day unknown and the gate refuses, for the reasonSpendBudgetgives: a ceiling that skips what it cannot price shows green while the real spend climbs. - What it does not do: it is off unless a limit is set, it never applies to the local backend (which spends electricity, not dollars), and the ceiling is per home rather than per key (the log does not record which key answered).
- Study 27, phase 3, second slice.
chimera/decisions/gate.py.
- What it does: a sliding 60-second window of requests and tokens. At 80% of a configured
Changed
-
glm-5.3-flashnow compacts at the context it was measured to read, 512k, instead of the 64k kept for unmeasured models.- The measurement:
bench/useful_context(RESULTS_glm53flash.md) ran 54 paired agent transcripts on the Novita route. It read 54/54 at 4k and 53/54 at 512k, with every rung within the registered margin. 512k was the top of the ladder, so the value is a lower bound. - What changes: the catalogue row carries
useful_k = 512. A surface that caps unmeasured models (the coding route) now budgets this model at 512,000 tokens rather than 64,000, and compacts near 409,000. - Adopted in its own change, as the pre-registration requires.
- The measurement:
-
A decision the spend/rate gate refused now raises an approval card instead of falling to the default.
- Before: a halt of any kind left the REVIEW band without a verdict and the kernel went on to its default, which is ALLOW. That is right for a model that was down (it said nothing about the action) and wrong for a meter that ran out.
- Now: when the gate refused the ask (
gate: "budget"or"rate"on the receipt), the band returns REVIEW withband: gateand a reason that says nothing judged the action, so a person is asked. An ordinary halt still goes on to the default, exactly as before, and a fixed rule still beats the band. - Who this touches: only a deployment that set
CHIMERA_DECISION_RPM,CHIMERA_DECISION_TPMorCHIMERA_DECISION_DAILY_USD(the gate does not exist without one). While the meter is out, each action no rule matched is a card. That was the owner's decision on 2026-09-29. - Study 27, phase 3, fourth slice.
-
After untrusted input, a write inside the workspace is a warning, not a card.
- Before: once a run had read anything external, every
write_file,edit_file,apply_patchandedit_batchasked for approval for the rest of the run.bench/injectionmeasured what that costs when nobody can answer: it refused every piece of legitimate work that began by reading something external. - Now, on the coding route (a screen is there to show a card to): those four tools go ahead,
the turn says so (
tainted_write, naming the tool, the path and where the taint came from), and the audit keeps ataint_write_warnedline. Still a card: anything that reaches a shell, a send over the network, a path outside the workspace, and the per-action check on what a write contains (a write of what the page said into a script or a config that something else will run). - What did not change: every other surface, every bench and the library default. The warning
is opt-in (
LedgeredTool(warn_workspace_writes=...)), so the narrowing thatbench/injectionand the authorization bench were measured with is the one they still run.
- Before: once a run had read anything external, every
-
A missing price and money spent are now warnings, not stops, until the person types a ceiling.
- Before: the desktop armed a US$1 ceiling on every turn, and a model with no known price stopped the run on its first call. Both read as the agent giving up for no reason the person chose.
- Now: the desktop sends no ceiling. A turn that has spent US$1 says so once (
spend_warn,warn_usd, default US$1 on the coding route), and a call on an unpriced model says that its spend is not being counted (price_unknown) and goes on. - What did not change: a ceiling someone typed is still a hard stop, checked before each call, and with a ceiling set an unpriced call still stops the run: a cap that skips what it cannot price shows green while the real spend climbs.
- Tests rewritten, with the reason:
SpendCeiling.test.tsx,Conversation.wave0.test.tsxandCode.spend.test.tsxasserted that a ceiling was armed by default. That premise is the decision that changed.
-
A repeating run is asked to change approach before it is cut.
- Before: the loop breaker warned at three identical calls and cut the run at four or five, and the warning was never shown to anyone.
- Now, where a person is waiting (the desktop and MCP coding route,
chimera chat,assist,tuiandagent): the first warning per tool adds a note asking the model to use different arguments or a different tool, or to say what is blocking it, and the breaker moves out to a net: ten identical calls, eight unchanged answers, six ping-pong cycles. A run that ignores the note is still stopped. - What did not change: the library default.
AgentConfig.loop_correctionis off, so every bench keeps the breaker it was measured with. The same commit turns the US$1 spend warning on for the terminal surfaces.
-
The step limit is a window, not a wall, where a person is waiting.
- Before: a turn stopped at its step limit and closed with a partial answer; the desktop had an opt-in "continue" toggle that sent the word "continue" up to three times, and no terminal surface had anything.
- Now (the desktop and MCP coding route,
chimera chat,assist,tuiandagent):AgentConfig.auto_continuemakesmax_stepsthe size of a window. At its end the run says so (steps_extended, with the step count) and takes another one, with no total ceiling. What ends the run is the work being done, a cancel, the loop breaker's net, a spend ceiling the person typed, or a full context. It no longer announces "2 steps left" when it is going on anyway. - What did not change: the library default.
auto_continueis off, somax_stepsis a wall for every bench, and the desktop's own toggle still exists and now has nothing left to do.
-
A conversation that outgrows its model compacts and goes on, and a model nobody measured is capped at 64k.
- Before: the compaction budget was 0.6 of the advertised window, so a model with a million
tokens and no measurement never compacted before whatever its real quality cliff is. The terminal
commands never set a budget at all, so an overflow there ended the run, and the desktop had no
badge for a run that could not be compacted any further (
context_stuck). - Now, where a person is waiting: the coding route and
chimera agent,chat,assistandtuicompact by default (0.6, or the fraction a client names), at the context the catalogue measured for the model (useful_k) and at 64k for a model it did not (AgentConfig.unmeasured_context_tokens). The cap is only a ceiling: a smaller window keeps its smaller number. The desktop sayscontext_stuckin ten languages. - What did not change: the library default. Both are off in
AgentConfig(), so every bench keeps the window-share budget it was measured with. The 64k is a precaution and not a measurement; the bench that would replace it (bench/useful_context, larger models) is pre-registered and waiting on a budget.
- Before: the compaction budget was 0.6 of the advertised window, so a model with a million
tokens and no measurement never compacted before whatever its real quality cliff is. The terminal
commands never set a budget at all, so an overflow there ended the run, and the desktop had no
badge for a run that could not be compacted any further (
-
The background-job limits (3 at once, 6 hours) are advice until the owner sets them.
- Before: a fourth job was refused, and a job still running at hour six was killed with everything it started. Both read as the agent giving up on something it was handed on purpose.
- Now: with
CHIMERA_JOBS_MAX_RUNNINGandCHIMERA_JOBS_MAX_RUNTIMEunset, the fourth job starts and its start message says it is past the usual three, and no job gets a deadline. Setting either one makes it a hard limit again, exactly as before: the cap refuses, and the runtime kills the tree and recordstimed_out. - What did not change: a registry built directly with numbers (a bench, a test) is a caller that meant them, so it stays hard.
-
A shell command that outlives its timeout keeps running as a job instead of being killed.
- Before:
run_shellkilled the command and its whole process tree when its timeout passed (60 seconds unless the model asked for more) and answered "timed out". A benchmark stage that takes tens of minutes died at the timeout the model happened to choose. - Now: where there is a job store and the host sandbox, the still-running command is adopted
as a background job. The answer says it is still running and was not stopped, and names
job_statusandjob_cancel. It can be cancelled like any job, its output is read from the log (written when it ends), and the next turn is told when it finishes. - What did not change: a
run_shellwith no job store, or on a container sandbox, still kills at the timeout, and so does one whose adoption a limit the owner set (CHIMERA_JOBS_MAX_RUNNING) forbids. The output of a still-running adopted command is not streamed:background=truefrom the start is still the way to watch one.
- Before:
-
Compaction on the default model now starts at the context it was measured to read.
- Before:
gpt-6-lunahad no measured useful context. The Code screen's budget was therefore 0.6 of its advertised 1,050k window, and it compacted at about 504k tokens. - Now:
bench/useful_contextheld an agent's context on 72 paired transcripts up to its 256k rung, at a median of 255,474 tokens, with every rung within −10 pp of 4k. The catalogue now carriesuseful_k = 255, a lower bound, so the budget is 255k and compaction lands at 204,000 tokens. - Cost of the run: US$ 4.67.
- Before:
Fixed
-
Two catalogue rows caught up with what the provider serves.
- What was wrong: the live check on main had been red since 2026-09-29.
deepseek-v4-flash-0731was priced 0.022/0.32 while the index quotes 0.0045/1.28, andllama-3.3-70b-instructpromised a 131k window its provider serves at 65,536. Both reach the user: the price prices a turn's receipt, and the window sets when a conversation compacts. - What happens now: both rows carry the live figures; the old price stays in
also_seen.
- What was wrong: the live check on main had been red since 2026-09-29.
-
A conversation deleted while one of its turns runs stays deleted.
- What was wrong: deleting a conversation (or a project's conversations) did not look at running turns. The turn went on and, when it finished, saved the conversation again, so a conversation the person had deleted came back.
- What happens now: deleting stops the turn with the same signal as Stop. That turn writes nothing more of the conversation: no transcript, no receipt, no history index entry.
-
An external agent in the middle of a turn is never closed to make room, nor for being idle.
- What was wrong: the ACP registry keeps at most four external agents (Claude Code, Gemini) alive, and closed the least recently used one when a fifth started, without asking whether it was mid-turn. The idle sweep had the same blind spot for a turn longer than an hour.
- What happens now: a turn is busy while it prompts, and only idle agents are closed. When every agent is busy, the registry goes over the limit for a while, and logs it, rather than killing work.
-
Dependencies past their open advisories. PyJWT 2.15.1 (15 advisories, one critical), urllib3 2.8.0 (three, two high) and brace-expansion 2.1.7 in the desktop tree. One PyJWT advisory has no fix yet, and glib still waits on a Tauri upgrade.
-
deepseek-v4-flash-0731carries the higher of the two prices it is quoted at. The route moved from 0.0045 to 0.0108 per million input tokens within a day; the row now reads 0.0108 so a fallback receipt never under-states what a turn cost, and 0.0045 stays as a price it has been seen at. -
The desktop updater's TLS library is patched.
- What was wrong:
rustls0.23.43, which the updater uses to download updates, had RUSTSEC-2026-0285 (TLS 1.3 handshake messages accepted across encryption level boundaries; medium). CI's Rust audit reported it from 2026-09-17 but is advisory, so the job stayed green and nobody saw it. - What happens now:
rustls0.23.45 andrustls-webpki0.103.15 (the version it needs); nothing else moved.
- What was wrong:
-
The status bar no longer mixes conversations, and it lists every turn running elsewhere, each with its own Stop.
- What was wrong: the bar reads one agent state, and every conversation wrote into it for as long as its handlers lived. Start a turn in A, switch to B and send, and when A finished it set "done, not busy", so B's Stop vanished while B was still running.
- Also wrong: the other running turns were invisible from the bar. With three working, it described one and offered one Stop.
- What happens now: a conversation publishes only while it is on screen, and hands the state back once when it leaves. Beside the bar's own subject, a chip reads "+N running". Its menu lists every other running coding turn (project · what was asked) from the server, each with a Stop that ends that turn on the server. It renders nothing while nothing else runs.
-
The Code screen comes back to the conversation you were in, per project.
- What was wrong: leaving the Code screen and coming back, or switching project and back, landed on a blank new conversation. The conversation was still in the list; the screen had forgotten it.
- What happens now: the last conversation of each project is remembered and reopened. New conversation forgets it, and so does deleting it from the sidebar.
-
A conversation holds at most four turns running or waiting. A share link reaches the guest route over the network, and nothing bounded how many turns a guest could pile onto one conversation, each a thread waiting on its lock. A fifth is refused with 429 before anything is built or announced.
-
A question waiting for a person says which project and conversation asked it.
- What was wrong: the status bar's list of waiting questions showed every conversation's questions in one dialog, and each said only what it would run and why. With two turns in two projects both asking to run a command, the wrong one could be approved.
- What happens now: the question file records the turn that asked (
run_id).GET /api/approvalsanswers each question with that turn's conversation and folder, read from the running turns (or from a background work's record). Each card in the list reads "From shop · Clean the build", or "From blog · background work “Resize images”". A question whose origin is unknown gets no line, never an invented one.
-
An undo offer and a finished job's news each stay with the conversation they came from.
- Undo offers: they lived in one list capped at 8 for the whole app. After eight editing turns anywhere, an older conversation's Undo button, and a finished background work's undo, answered "nothing to undo". Each conversation now keeps its 20 most recent offers, and the app keeps 200 in all, oldest first out. Each offer holds only the files its turn changed, not a copy of the whole folder.
- Finished jobs: "a background job finished" went to the next turn of ANY project. That turn marked it reported and could not read its output, because the job tools are fenced to the turn's folder, so the project that started the job never heard. A turn is now told only about jobs that ran inside its own folder.
-
Undo takes back what its own turn changed, and one conversation at a time edits a folder.
- What was wrong: undo restored the whole folder to the snapshot taken before the turn. With two conversations in one folder, undoing one also reverted the other's edits and anything the person had typed since; outside a git repository it deleted the other's new files. Nothing kept two turns from editing the same folder at once, either, so a snapshot, a verification or an undo could describe a mix of both.
- What undo does now: a turn records what it changed: the difference between the folder before it and right after it, so edits made through the shell count too. Undo puts back only those files. A file that changed again after the turn is left as it is, and the screen names it ("except N file(s) that changed after this turn"). The rules against deleting inside a git repository or after a truncated snapshot still apply.
- What turns in one folder do now: they take turns, background works included, as works already did among themselves. A turn that waits is told why. Stop still reaches it while it waits, and it never starts. Different folders still run at once.
-
A long turn followed late comes back whole, and idle conversations stop holding frames.
- What was wrong: each conversation kept its last 4000 live frames, and a coding turn streams a frame per token. A screen that came back to a long turn got its tail: no opening frame to draw the row on, and the answer without its start. Every conversation opened in the process also kept its frames until the app closed.
- What happens now: a replay that asks for frames the buffer dropped gets, for each turn still in it, the opening and the missing frames back from the run log, which records the session number of every frame. Conversations nobody watches and nothing runs in for 30 minutes drop their frames (the numbering goes on), and a deleted conversation's are dropped.
-
A memory write no longer fails on Windows when the file is busy for longer than 0.2 s. The antivirus or a paused reader can hold a file open past the old retry window, and the write then failed, losing the memory being saved. Retries now back off from 10 ms to 250 ms, about 3.8 s in all.
-
Two turns on the same conversation both stay in it.
- What was wrong: the coding route loaded the conversation before taking its lock and saved it inside the lock. A second turn on the same conversation (the owner and a guest on a shared link, or two tabs) waited for the first, then ran on the history it had loaded before the first finished, and its save erased the first turn's exchange. The comment above the lock promised the protection the code did not give.
- What happens now: every use of the lock reads the conversation again first. The turn's receipt, written after a verification that can take minutes, goes on top of what is stored, never back over a later turn with an older copy. A spoken work's note in its parent conversation does the same.
- Saves are atomic: a temporary file, then a rename. A reader that met half a file used to read the
conversation as unreadable, and
loadstarts an unreadable conversation fresh.
-
Stop ends a coding turn on the server, not only the screen's view of it.
- What was wrong: the Stop button aborted the screen's request and nothing else. There was no route to stop a coding turn, and the stop signal reached the agent loop only for background works. A turn the person had stopped went on calling the model, editing files and spending until it ended by itself, while the screen said it had stopped. Found reading the code while mapping how several conversations run at once.
- What happens now:
POST /api/code/turns/{turn_id}/stopraises the turn's stop signal. Chimera's own loop checks it once per step, so the step in progress finishes first. An external agent's prompt (Claude Code, Gemini over ACP) is cancelled at once. A turn that is not running answers 404, never a stop that reached nothing. - Where it applies: the Stop button uses it for a turn the screen started and for one it was following after coming back to the conversation.
-
"Let the agent try to fix it" says why it cannot start. With a run already working, the button handed the fix to a session that refused it without a word, and the click did nothing. It is now disabled, with a line saying why.
-
External links open in the system browser in the desktop app.
- What was wrong: the links that open a new window (the repository, the releases, the MCP and skill catalogues) did nothing in the desktop app. The runtime refuses every new window unless told otherwise, and it said nothing. The same links worked in the browser build.
- What happens now: the native side hands such a link to the operating system, and the page itself still opens nothing.
- What is refused: only an
httporhttpsaddress with a host goes through. It must carry no credentials and be off the app's own origin.file:,javascript:,mailto:and custom protocols are refused, so a link cannot make the system run a program. - New dependency: the
opencrate, 5.4.4. It is what Tauri's own opener plugin uses underneath, used here from Rust only; the plugin would add webview commands this app does not grant.
-
The local decider no longer judges a truncated state. Ollama keeps only half of its context window for the prompt and cuts the rest with no error. The default window here was 4,096 tokens, so the local System One model read at most ~2,050 tokens (about 11,000 characters of prose), while the verified-answers check sends attached sources of up to 14,000 characters. Every call now asks for a 16,384-token window (8,192 tokens of prompt), and a prompt that still fills it is a halt: the answer ships unverified, never judged on text the model did not see. Measured before the change on this backend (
bench/jevbench_local): 231 of 231 answers identical with the larger window. -
One Chimera Desktop server per data folder.
- What was wrong: a second copy of the app (a second click on the icon, or
chimera appin a terminal) could serve the same data folder. Each server keeps in memory the turns it runs and the folders they edit, so two of them each believed they were alone: the one-writer-per-folder rule held only inside each, and Stop in one window could not reach a turn the other ran. - What happens now: the server claims its data folder before building anything, with a lock the OS drops when the process ends. A second server on the same folder says where the first one is and exits with code 3.
- What was wrong: a second copy of the app (a second click on the icon, or
-
chimera decideexits 1 when a question fails. It exited 0 even when the backend was down or the state overflowed, so a CI step or a script could not tell "answered" from "failed". Now: 0 every question answered; 1 at least one failed, after the JSON is printed (or every JSONL line written) in full; 2 usage, or a question the linter refuses. -
An autonomous run waits for its folder the way a coding turn does.
- What was wrong: turns took one lock per folder; runs took none, so another window or client could start a run beside a turn, or a second run, in the same folder.
- What happens now: turns and runs share one lock per folder. A run whose folder is busy says so on its feed and waits; Stop still reaches it while it waits.
-
apply_patchno longer writes a stray conflict marker into a file.- What went wrong: a hunk with a second
=======line put that line into the replacement text. The marker landed in the source file, and the tool still answered "applied". This was seen live, with the desktop agent editing through the app. - Now: a hunk whose SEARCH or REPLACE holds a marker line is refused, the file is left
untouched, and the error points to
edit_filefor a file that really contains such a line.
- What went wrong: a hunk with a second
-
Portuguese advice and appraisal requests on attached documents are no longer checked as if they were questions.
- What went wrong: on a fresh set written by another model family, the grounded-question classifier read 5 of 80 tasks as questions. Four were Portuguese, and their English twins passed. Examples: "Será que dá pra reescrever…", "O que você mudaria…", "Me ajuda a decidir…", "parece adequada… Justifique".
- The fix: these frames now read as tasks, and courtesy questions ("você poderia me dizer…") are still questions.
- The stake is small. Forced through the check, 60 tasks lost one real attempt
(
bench/grounded_task_declines).