8 Downloads Updated 2 days ago
ollama run jrlabs01/s1
Updated 2 days ago
2 days ago
3fd58603b806 · 18GB
s1 makes typed decisions over text or JSON (pick an option, answer yes/no, or place something on an ordered scale) and gives a calibrated probability for every option, from one forward pass. This is the int4 build of j-raghavan/s1-gemma4-26b-decision, a fine-tune of Gemma 4 26B-A4B. Apache-2.0.
Accuracy of this build (held-out test sets): JevBench subset 0.801 (bf16: 0.808), structured decisions 0.966 (bf16: 0.971); neither difference is significant. About 26 GB of memory while loaded.
s1 reads the probabilities of the option letters after the prefix Answer:. Send raw prompts that start with
<bos>; Ollama’s chat formatting changes the prompt and the answers.
The easiest way is the repository’s HTTP API, which does this for you and adds calibration:
ollama pull jrlabs01/s1
git clone https://github.com/j-raghavan/s1-decision-model && cd s1-decision-model
S1_MODEL=jrlabs01/s1 uv run --extra api uvicorn api.server:app --port 8000
curl -s localhost:8000/v1/decisions -H 'content-type: application/json' -d '{
"state": {"ticket": "I was charged twice for order 4471."},
"questions": {"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Payments and refunds", "shipping": "Deliveries", "tech": "Bugs"}}}}'
Calling Ollama directly: POST /api/generate with "raw": true, "logprobs": true, "top_logprobs": 20,
"options": {"num_predict": 1, "temperature": 0} and the prompt format in
examples/quickstart.py, prefixed
with <bos>.
Benchmarks, methodology and training data: github.com/j-raghavan/s1-decision-model. s1 is an independent project, not affiliated with TypeSafe AI (Jev).