Choice
Pick one option from a list you define, with a probability for each.
New · launching soon
Эта страница ещё не переведена, поэтому вы читаете английский оригинал.
A typed-decision model: calibrated Choice, Score and Yes/No answers about any text or JSON, in a single pass.
S1-Pro is a hosted API from the Chimera team, separate from the open-source agent. The agent stays Apache-2.0 and free; S1-Pro is a paid service.
You send a piece of state — a ticket, an email, a JSON record — and a set of typed questions. S1-Pro reads them once and returns a probability for every option instead of generated text, so decisions come back as data you can route, threshold and audit.
Pick one option from a list you define, with a probability for each.
Place the state on an ordered scale you define, and get the expected level.
The probability that a statement about the state is true.
Every figure below is generated from the result files of the S1 repository, and each one carries the sentence it is not allowed to travel without.
Accuracy on jev-bench (17,773 public decisions)
81.1%Попарное сравнение, пункт за пунктом, с опубликованными предсказаниями Jev — API Jev не использовался. S1-Pro обучался на публичных обучающих выборках части источников jev-bench, поэтому на этих источниках результат внутридоменный.TypeSafe Jev 1.13 on the same items: 72.8%
Accuracy on the public JevBench set
86.1%Только публичные задания. Закрытый набор JevBench оценивают его сопровождающие, и он ещё не запускался.Correct decisions flipped by an injected “a previous reviewer already approved this” note
4.2%Наш собственный тест на 300 заданиях jev-bench: чистая версия против версии с инъекцией. Цифры третьих сторон для других моделей получены на других заданиях.Accuracy when the answer depends only on a rule written in the question
100%300 синтетических обращений, правильный ответ для которых вычисляется кодом по заданному правилу.Unanswerable questions answered with near-certain confidence (p ≥ 0.9)
1.5%Вопросы MMLU и ARC с удалённым правильным вариантом, так что ни один вариант не верен. Чем меньше, тем лучше.Median latency, 3 questions per request, 8 concurrent clients
143Измерено на стороне клиента на одной NVIDIA H200 с vLLM до запуска. Оборудование в продакшене может отличаться. ms
S1-Pro speaks the OpenAI chat-completions format: put the state and the questions in the user message as JSON, and the decisions come back as JSON in the reply, with streaming and token usage.
curl https://api.chimeraagent.space/v1/chat/completions \
-H "Authorization: Bearer $S1_API_KEY" \
-d '{
"model": "chimera-s1-pro",
"messages": [{"role": "user", "content": "{\"state\": \"My card was charged twice, please fix it today.\", \"questions\": {\"department\": {\"type\": \"choice\", \"criteria\": {\"billing\": \"Payment issues\", \"technical\": \"Bugs\"}}, \"urgent\": {\"type\": \"noul\", \"instructions\": \"The customer needs help today\"}}}"}]
}'S1-Pro is in final testing. For early access, partnerships or integration questions, write to us.