Ir al contenido

Blog

Releasesv0.15.0

v0.15.0 — Trustworthy Verifier & Write Security

Esta es la nota de release tal como fue publicada, no una reescritura de ella. Las notas de release se publican en el idioma en que fueron escritas.

The M18 cycle, grounded in an 8-category arXiv sweep — five Tier-1 items that harden the weakest link in the "prove a weak model performs like a frontier one" thesis (the verifier), plus write-side security and a measured skill-lifecycle loop. Every item is an integration (never a reimplementation), opt-in, degrades safely, and is fake-tested. The headline shipped with a measured paired A/B and an honest retraction.

Added

  • solve --gen-tests — spec-grounded test generation (M18-1). With no --verify command, the fitness gate falls back to an LLM judging coverage — a proxy that rubber-stamps wrong code. --gen-tests generates executable pytest grounded in the task's atomic requirements and runs it as the gate. Measured (paired A/B, goldilocks mistral-small-24b, n=6): gen-tests 3/6 vs 0/6 for the coverage grade (Δ +50pp, 3-0 discordant, CI [-6.1%,+50%] — not significant at n=6, reported as-is). Honest caveat: the win is a resolve-rate win, not the false-positive reduction I pre-registered — that prediction was retracted.
  • Cross-provider, decomposed envelope auditing (M18-2). The spot-check auditor grades three named failure classes separately (invented / dropped / contradiction) and can run on a distinct model (orchestrate --verify-model) so a model doesn't grade its own family's output.
  • Declared write-region for the file-writers (M18-3). solve --write-region "src/**,*.py" makes write_file/edit_file/apply_patch refuse a write outside the declared globs — blocking an injected instruction from rewriting an unrelated file.
  • Measured skill-lifecycle loop (M18-4). A provisional tier: new skills run on probation, then chimera skills-lifecycle [--apply] auto-promotes the proven ones and demotes regressions — from measured stats, never self-report.
  • PROBE — best-arm identification with a cheap-proxy control variate (M18-5). chimera probe-select picks the best model/config with fewer expensive grades the better a cheap proxy correlates, and stays unbiased when it doesn't.

Gate: ruff clean, mypy clean, 1250+ tests. Still alpha.

PyPI: `pip install chimera-agent==0.15.0`

Leer la release en GitHub