Перейти к содержимому

Все команды

chimera skillcard-bench

A/B reasoning with vs without injected TRS skill cards. Calls real models.

Параметры

  • --tasksstrпо умолчанию: 'hard'

    Task suite: hard | big | demo. 'big' = 24 traps for a tighter paired CI.

  • --kintпо умолчанию: 1

    How many cards to retrieve per task.

  • --min-overlapintпо умолчанию: 2

    Relevance gate: inject a card only on >= N shared query terms (0=off).

  • --max-linesintпо умолчанию: 3

    Render budget: max lines per injected card.

  • --use-storeboolean

    Bench your own learned cards (skills.json) instead of the demo set.