48 4 days ago

Decision Model: Answers yes/no, multiple-choice and scoring questions about any text, returning a probability for every option. More accurate than nimble:9b on JevBench.

vision tools decision thinking
curl http://localhost:11434/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "model": "aminroudaki/decisio:q8_0",
    "state": "Hello World",
    "questions": {
      "says_hello": {
        "type": "noul",
        "instructions": "Does the state text contain a greeting?",
        "criteria": {
          "true": "The state text contains a greeting.",
          "false": "The state text does not contain a greeting."
        }
      }
    }
  }'

Details

4 days ago

29bc7aac46fd · 39GB

qwen35moe
·
35.5B
·
Q8_0
clip
·
447M
·
F16
Answer with the letter only.
{ "num_ctx": 8192 }

Readme

Decisio

decisio answers closed questions about any text, with a probability for every option. Give it a piece of text (a support ticket, a document, a chat message, a JSON record) and one or more questions; it returns, for each question, the chosen answer and how likely each option is.

  • Yes/no: “Is this message toxic?”, “Does this ticket need a reply within the hour?” You get the probability of yes.
  • Choice: route, classify or pick among up to 26 options: “Which team should handle this?”, “Which answer is correct?” You get the top option and every option’s probability.
  • Score: rate on a scale you define: “How severe is the impact, 0 to 3?” You get the expected level and each level’s probability.
  • Many questions at once: up to 64 questions about the same text in one request.
  • Fast: one forward pass per question, no text generation. On a laptop (Apple M5 Pro) a decision takes about 0.24 s at the median.

Measured 2026-09-30.

Requires Ollama 0.35.1 or later: the decision route (POST /v1/systemone) serves this model through its declared decision capability.

How to use it

Requires Ollama 0.35.0 or later.

ollama pull aminroudaki/decisio

Send the text as state and your questions to Ollama’s decision endpoint:

curl http://localhost:11434/v1/systemone -d '{
  "model": "aminroudaki/decisio",
  "state": "Hi, since this morning none of our 40 staff can log in to the dashboard. We get \"session expired\" right after entering the password. Payroll is due today.",
  "questions": {
    "urgent": {
      "type": "noul",
      "instructions": "Does this ticket need a response within the hour?"
    },
    "category": {
      "type": "choice",
      "instructions": "Which team should handle this ticket?",
      "criteria": {
        "billing": "Invoices, payments, refunds",
        "access": "Login, passwords, permissions",
        "bug": "Something in the product behaves wrongly",
        "other": null
      }
    },
    "impact": {
      "type": "score",
      "instructions": "Rate the business impact using only the reported facts.",
      "criteria": [
        "No function impaired",
        "One user impaired, with a workaround",
        "Many users blocked from a core function",
        "Data loss or legal exposure"
      ]
    }
  }
}'

Response (abridged, probabilities rounded):

{
  "answers": {
    "urgent":   {"type": "noul", "noul": 0.965},
    "category": {"type": "choice", "choice": "access",
                 "probabilities": {"billing": 0.000, "access": 0.999, "bug": 0.000, "other": 0.000}},
    "impact":   {"type": "score", "score": 2.000,
                 "probabilities": {"0": 0.000, "1": 0.000, "2": 0.999, "3": 0.000}}
  },
  "usage": {"input_tokens": 849, "output_tokens": 4}
}
  • noul (yes/no): noul is the probability of yes. Optional criteria {"true": "...", "false": "..."} describe what yes and no mean.
  • choice: criteria maps each option key to a description, or null when the key speaks for itself. Up to 26 options.
  • score: criteria lists the levels in order, from 0. score is the expected level.
  • state can be plain text or a JSON object or array.

Which tag:

Tag Download Use
latest = q4_k_m 22 GB Recommended; as accurate as q8_0 on every measure below, and faster
q8_0 38 GB Higher-precision weights, if you have the memory

How it performs

Measured on two test sets: - 1,250 held-out items from seven public tasks: BoolQ, MMLU, two MMLU-Pro sets, SciFact twice, ToxicChat; - the 231 published JevBench items (easy, standard, hard), run with the benchmark’s own harness.

decisio, nimble:9b and tev1 were run through Ollama’s /v1/systemone on the same items. Jev 1.13.0’s row is the JevBench operator’s published result on the same 231 JevBench items (results/v1.2/jevbench-v1.2-per-task.json in github.com/fstandhartinger/jevbench, JevBench v1.3.0). It publishes right or wrong per item, so there is no calibration or suite figure for Jev.

System Suite accuracy Suite ECE decisio q4_k_m - system, suite JevBench accuracy JevBench ECE decisio q4_k_m - system, JevBench
decisio q4_k_m 0.775 0.089 - 0.853 0.045 -
decisio q8_0 0.770 0.099 +0.6 [-0.4, +1.6] 0.848 0.054 +0.4 [-1.7, +2.6]
nimble:9b 0.750 0.069 +2.6 [+0.5, +4.6] 0.797 0.097 +5.6 [+0.9, +10.4]
tev1:4b 0.678 0.061 +9.8 [+7.4, +12.0] 0.762 0.094 +9.1 [+3.9, +14.3]
tev1:0.8b 0.462 0.151 +31.3 [+28.2, +34.5] 0.615 0.136 +23.8 [+17.3, +30.3]
Jev 1.13.0 (JevBench operator’s results) - - - 0.866 - -1.3 [-4.8, +2.2]

Bold marks where decisio leads with the interval clear of zero.

Differences are in points of accuracy, with paired 95% intervals, over the same 1,250 and 231 items for every system (Jev: the 231 JevBench items).

  • The most accurate of the models tested in Ollama. decisio is ahead of each of them on both sets, with intervals clear of zero (q8_0 is decisio at higher precision, and level):
    • nimble:9b by 2.6 and 5.6 points;
    • tev1:4b by about 9 to 10 points, and 18 on the JevBench hard tier;
    • tev1:0.8b by 24 to 31.
  • Against Jev 1.13.0, on the JevBench operator’s published results for the same 231 items: level, -1.3 [-4.8, +2.2]; hard tier 0.721 against 0.730.
  • tev1 needs a longer context for long inputs. It was measured with num_ctx raised to 8192. At its published default (2,050 tokens) Ollama refuses 36 of the 111 hard JevBench items and 2 of the suite items as too long; on the other suite items its top answer is the same in 1,247 of 1,248.
  • Calibration: on JevBench, decisio’s probabilities are the best calibrated of the Ollama models here, better than nimble:9b’s and tev1’s with intervals clear of zero. On the suite, tev1:4b’s are better calibrated (ECE 0.061 against 0.089), and nimble:9b’s (0.069) are level within the interval.

decisio against the Decisio server, with Brier scores and latency:

Measure decisio q4_k_m decisio q8_0 nimble:9b Decisio server
Suite accuracy (1,250) 0.775 0.770 0.750 0.765
Suite ECE / Brier 0.089 / 0.335 0.099 / 0.332 0.069 / 0.340 0.025 / 0.322
JevBench accuracy (231) 0.853 0.848 0.797 0.844
JevBench hard tier (111) 0.721 0.721 0.622 0.694
JevBench ECE / Brier 0.045 / 0.210 0.054 / 0.215 0.097 / 0.289 0.040 / 0.222
Latency p50 / p95, M5 Pro laptop 241 / 488 ms 291 / 606 ms 345 / 657 ms not comparable (GPU server)
  • Accuracy: as accurate as the Decisio server: +1.0 [-0.8, +2.8] on the suite and +0.9 [-3.0, +4.8] on JevBench for q4_k_m.
  • Calibration is weaker on exam-style questions. Ollama applies no temperature, so suite ECE is 0.089-0.099 against 0.025 for the Decisio server (MMLU-Pro is the worst task). On JevBench, calibration is level with the Decisio server.
  • How to read the numbers:
    • ECE is expected calibration error (10 equal-mass bins, top answer); Brier is multiclass. Lower is better for both.
    • Intervals are paired bootstrap, 4,000 draws.
    • Nothing was tuned on any evaluated item.
    • Latency: one question per request, model loaded, on one Apple M5 Pro (64 GB); the Mac was not fully idle for the decisio passes.

The Decisio server

https://github.com/aminry/decisio

The Decisio server adds fitted calibration, large taxonomies, task registration from labelled examples and image input.

Limits

  • Options: 2 to 26 per question (Ollama’s limit). Larger taxonomies are refused; our 77-intent Banking77 test set could not be run.
  • Request size: 64 KiB per request, at most 64 questions.
  • Context: 8,192 tokens, about twice the longest item tested (4,014 tokens). Longer texts need a larger num_ctx.
  • Uncalibrated probabilities: they come straight from the model. Treat them as a ranking plus a rough confidence, and recalibrate on your own data if you threshold on them.
  • Run-to-run variation: probabilities on the same item can differ between runs by up to about 0.1, because Ollama reuses cached prompt parts from earlier requests. Top answers were unchanged on every spot-checked item. If you threshold on probabilities, measure on your own traffic.
  • Ollama builds the prompt itself, so the Decisio server’s published numbers do not apply here; the numbers on this page do. On about one question in seven, decisio’s top answer differs from the Decisio server’s, and on multiple-choice knowledge questions about one in four; the accuracies are level.
  • English only tested. Items are public English benchmarks. No safety evaluation beyond ToxicChat classification.

Licence

Apache-2.0. Based on Qwen3.6-35B-A3B; GGUF quantisations by bartowski.