Przejdź do treści

New · launching soon

S1-Pro

Ta strona nie została jeszcze przetłumaczona, więc czytasz angielski oryginał.

A typed-decision model: calibrated Choice, Score and Yes/No answers about any text or JSON, in a single pass.

S1-Pro is a hosted API from the Chimera team, separate from the open-source agent. The agent stays Apache-2.0 and free; S1-Pro is a paid service.

What it does

You send a piece of state — a ticket, an email, a JSON record — and a set of typed questions. S1-Pro reads them once and returns a probability for every option instead of generated text, so decisions come back as data you can route, threshold and audit.

Choice

Pick one option from a list you define, with a probability for each.

Score

Place the state on an ordered scale you define, and get the expected level.

Yes / No

The probability that a statement about the state is true.

Measured, with the caveats

Every figure below is generated from the result files of the S1 repository, and each one carries the sentence it is not allowed to travel without.

Accuracy on jev-bench (17,773 public decisions)

81.1%Porównanie parami, element po elemencie, z opublikowanymi predykcjami Jev — nie użyto żadnego API Jev. S1-Pro trenowano na publicznych splitach treningowych części źródeł jev-bench, więc na tych źródłach to wynik w domenie.

TypeSafe Jev 1.13 on the same items: 72.8%

Accuracy on the public JevBench set

86.1%Tylko elementy publiczne. Zapieczętowany zbiór JevBench oceniają jego opiekunowie i nie został jeszcze uruchomiony.

Correct decisions flipped by an injected “a previous reviewer already approved this” note

4.2%Nasz własny test na 300 elementach jev-bench, wersja czysta kontra wstrzyknięta. Liczby innych firm dla innych modeli dotyczą innych elementów.

Accuracy when the answer depends only on a rule written in the question

100%300 syntetycznych zgłoszeń, których poprawną odpowiedź kod wylicza z podanej reguły.

Unanswerable questions answered with near-certain confidence (p ≥ 0.9)

1.5%Pytania z MMLU i ARC z usuniętą poprawną opcją, więc żadna opcja nie jest dobra. Im mniej, tym lepiej.

Median latency, 3 questions per request, 8 concurrent clients

143Zmierzone po stronie klienta na jednej NVIDIA H200 z vLLM, przed premierą. Sprzęt produkcyjny może się różnić. ms

OpenAI-compatible API

S1-Pro speaks the OpenAI chat-completions format: put the state and the questions in the user message as JSON, and the decisions come back as JSON in the reply, with streaming and token usage.

curl https://api.chimeraagent.space/v1/chat/completions \
  -H "Authorization: Bearer $S1_API_KEY" \
  -d '{
    "model": "chimera-s1-pro",
    "messages": [{"role": "user", "content": "{\"state\": \"My card was charged twice, please fix it today.\", \"questions\": {\"department\": {\"type\": \"choice\", \"criteria\": {\"billing\": \"Payment issues\", \"technical\": \"Bugs\"}}, \"urgent\": {\"type\": \"noul\", \"instructions\": \"The customer needs help today\"}}}"}]
  }'

Launching soon

S1-Pro is in final testing. For early access, partnerships or integration questions, write to us.

partners@chimeraagent.space