3 19 hours ago

Fast 2B decision model for System One. Picks options, gives yes/no probabilities and scores in one forward pass. Tuned on games and general decisions, based on MiniCPM5-2B.

decision 2b
curl http://localhost:11434/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MrScratchcat22/aurelis:2b",
    "state": "Hello World",
    "questions": {
      "says_hello": {
        "type": "noul",
        "instructions": "Does the state text contain a greeting?",
        "criteria": {
          "true": "The state text contains a greeting.",
          "false": "The state text does not contain a greeting."
        }
      }
    }
  }'

Details

19 hours ago

c297924173f6 · 2.7GB

llama
·
2.52B
·
Q8_0
<s><|im_start|>user {{ range .Messages }}{{ .Content }}{{ end }}<|im_end|> <|im_start|>assistant <th
{ "num_ctx": 4096 }

Readme


license: apache-2.0 base_model: openbmb/MiniCPM5-2B base_model_relation: finetune library_name: transformers pipeline_tag: zero-shot-classification language: - en datasets: - tasksource/tasksource-jev-typed-decisions tags: - decision-model - system-one - ollama - gguf

- minicpm

Aurelis 2B

Aurelis is a small decision model. You give it a situation and one or more questions with fixed options, and it returns a probability for every option from a single forward pass. It never generates text, which keeps it fast: a whole set of questions is answered in the time a chat model needs for a few tokens.

It speaks Ollama’s System One API (/v1/systemone) with the tev1 prompt encoding, so it runs in Ollama the same way Tev1, Clef and Nimble do. It is fine-tuned from MiniCPM5-2B on game decisions and general human-labelled decisions.

Run it with Ollama

Needs Ollama 0.35 or newer.

ollama pull mrscratchcat22/aurelis:2b
curl http://localhost:11434/v1/systemone -d '{
  "model": "mrscratchcat22/aurelis:2b",
  "state": "Blackjack. Player has 10 and 6 (hard 16). Dealer shows a 10. No surrender.",
  "questions": {
    "move": {"type": "choice", "instructions": "Best move for the player?",
             "criteria": {"hit": "take another card", "stand": "keep the hand"}},
    "bust": {"type": "noul", "instructions": "Will the player bust if they take one more card?"}
  }
}'

There are three question types. choice picks one of 2 to 24 options and returns the pick, a probability per option and a confidence. noul returns the probability that a statement is true. score rates on an ordered scale of 2 to 24 levels and returns the expected level. One request takes up to 64 questions about the same state. Text only.

Run it with Transformers

The model reads the exact JSON message Ollama sends and answers with the letter of an option. Reading the logits of those letters gives the probabilities. The calibration is built into the weights, so a plain softmax is already calibrated.

import json, torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "MrScratchcat/aurelis-2b"
device = "cuda" if torch.cuda.is_available() else "cpu"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16).to(device).eval()

def decide(state, question, options):
    """options: {key: description}. Returns {key: probability}."""
    opts = [{"label": chr(65 + i), "key": k, "description": d or k} for i, (k, d) in enumerate(options.items())]
    msg = json.dumps({"state": state, "question": question, "options": opts}, ensure_ascii=False)
    text = tok.apply_chat_template([{"role": "user", "content": msg}], add_generation_prompt=True,
                                   tokenize=False, enable_thinking=False)
    ids = tok(text, add_special_tokens=False, return_tensors="pt").to(device)
    with torch.no_grad():
        logits = model(**ids, logits_to_keep=1).logits[0, -1]
    letters = [tok(o["label"], add_special_tokens=False)["input_ids"][0] for o in opts]
    return dict(zip(options, torch.softmax(logits[letters].float(), -1).tolist()))

state = "Blackjack. Player has 10 and 6 (hard 16). Dealer shows a 10. No surrender."
print(decide(state, "Best move for the player?", {"hit": "take another card", "stand": "keep the hand"}))
print(decide(state, "Will the player bust if they take one more card?", {"false": "No", "true": "Yes"})["true"])

For a yes/no question, use the options false: No and true: Yes in that order, as Ollama does. For a scale, use the level numbers 0, 1, 2, … as keys.

Evaluation

Held-out rows the model never saw in training, scored with the prompts Ollama sends. Chance is the score of picking an option at random. Brier and ECE measure how trustworthy the probabilities are; lower is better.

set rows accuracy chance Brier ECE
general decisions 1500 68.1% 35.7% 0.127 0.032
chess 128 75.0% 12.0% 0.042 0.099
poker 127 84.3% 42.5% 0.087 0.046
blackjack 109 96.3% 36.6% 0.017 0.061
grid paths 101 93.1% 55.5% 0.068 0.212
tic-tac-toe 86 66.3% 44.6% 0.123 0.196
Connect Four 66 45.5% 18.4% 0.096 0.116
Doom (rule-labelled) 139 98.6% 30.3% 0.006 0.008
FNAF (rule-labelled) 76 100.0% 19.3% 0.000 0.004
bets 8 75.0% 50.0% 0.217 0.144
all 2340 73.9% 34.7% 0.100 0.033

The general decisions come from the validation split of tasksource-jev-typed-decisions: classification, NLI, QA, sentiment and similar tasks with human labels.

These are the full bf16 weights. The Q8_0 build on Ollama scores 72.6% overall on the same rows (ECE 0.042).

Speed, measured with a plain Transformers server in bf16 on an RTX 4070 Laptop GPU (median of 20 runs): 88 ms for one question, 116 ms for four, 447 ms for sixteen.

Training

LoRA fine-tune of MiniCPM5-2B, merged into the weights afterwards. Rank 16 (alpha 32, dropout 0.05) on every linear layer, learning rate 1e-4 with a cosine schedule, batch 16, one epoch over about 55,000 examples, bf16, about five hours on an 8 GB RTX 4070 Laptop GPU. The loss is cross-entropy between the target distribution and the softmax over the allowed answer letters only; nothing else in the vocabulary is trained.

The data has two parts. 40,000 general decisions come from tasksource-jev-typed-decisions, with soft targets taken from annotator votes where the source has them. 15,360 game decisions were generated and labelled by program: chess positions scored by Stockfish at depth 13, poker spots by exact equity, blackjack hands by exact expected value (infinite deck, dealer stands on soft 17, double after split), simple bets by expected value, tic-tac-toe by minimax, Connect Four by win and block detection, grid movement by shortest path, and Doom and FNAF situations by hand-written rules. No game files were used.

After training, a temperature of 1.1 was fitted on the held-out rows to calibrate the probabilities, then folded into the output layer.

Limitations

The Doom and FNAF labels come from hand-written rules, so the near-perfect scores mean the model learned those rules, not that it plays those games well. Connect Four and tic-tac-toe are its weakest games, and its confidence there runs higher than its accuracy. Bets has too few evaluation rows to say much.

This first version was trained on a plain-text question layout, where it scores 76.0% on the same rows. Ollama sends JSON, where it scores 73.9%; the table above uses Ollama’s format.

The general data comes from many sources with their own licenses, and each row of tasksource-jev-typed-decisions names its source and license. Some of those sources do not allow commercial use, so check them before using Aurelis commercially. The base model, MiniCPM5-2B, is Apache-2.0.