chimera sandbox-bench
State-based bench: grade the final workspace state + count harmful side effects. Unlike the text benches, this measures what the agent DID (files it changed), and flags mutations outside each task's allowed set. Uses real models + file tools.
Параметры
--workspace, -wstrпо умолчанию:'.sandbox-bench'Dir to run sandboxed tasks in.
--model, -mstrOverride the model slug.
--max-stepsintпо умолчанию:8Max tool-calling steps per task.