3 Downloads Updated 17 hours ago
curl http://localhost:11434/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"model": "MrScratchcat22/aurelis:2b",
"state": "Hello World",
"questions": {
"says_hello": {
"type": "noul",
"instructions": "Does the state text contain a greeting?",
"criteria": {
"true": "The state text contains a greeting.",
"false": "The state text does not contain a greeting."
}
}
}
}'
license: apache-2.0 base_model: openbmb/MiniCPM5-2B base_model_relation: finetune library_name: transformers pipeline_tag: zero-shot-classification language: - en datasets: - tasksource/tasksource-jev-typed-decisions tags: - decision-model - system-one - ollama - gguf
Aurelis is a small decision model. You give it a situation and one or more questions with fixed options, and it returns a probability for every option from a single forward pass. It never generates text, which keeps it fast: a whole set of questions is answered in the time a chat model needs for a few tokens.
It speaks Ollama’s System One API (/v1/systemone) with the tev1 prompt encoding, so it runs in Ollama the same way Tev1, Clef and Nimble do. It is fine-tuned from MiniCPM5-2B on game decisions and general human-labelled decisions.
Needs Ollama 0.35 or newer.
ollama pull mrscratchcat22/aurelis:2b
curl http://localhost:11434/v1/systemone -d '{
"model": "mrscratchcat22/aurelis:2b",
"state": "Blackjack. Player has 10 and 6 (hard 16). Dealer shows a 10. No surrender.",
"questions": {
"move": {"type": "choice", "instructions": "Best move for the player?",
"criteria": {"hit": "take another card", "stand": "keep the hand"}},
"bust": {"type": "noul", "instructions": "Will the player bust if they take one more card?"}
}
}'
There are three question types. choice picks one of 2 to 24 options and returns the pick, a probability per option and a confidence. noul returns the probability that a statement is true. score rates on an ordered scale of 2 to 24 levels and returns the expected level. One request takes up to 64 questions about the same state. Text only.
The model reads the exact JSON message Ollama sends and answers with the letter of an option. Reading the logits of those letters gives the probabilities. The calibration is built into the weights, so a plain softmax is already calibrated.
import json, torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "MrScratchcat/aurelis-2b"
device = "cuda" if torch.cuda.is_available() else "cpu"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16).to(device).eval()
def decide(state, question, options):
"""options: {key: description}. Returns {key: probability}."""
opts = [{"label": chr(65 + i), "key": k, "description": d or k} for i, (k, d) in enumerate(options.items())]
msg = json.dumps({"state": state, "question": question, "options": opts}, ensure_ascii=False)
text = tok.apply_chat_template([{"role": "user", "content": msg}], add_generation_prompt=True,
tokenize=False, enable_thinking=False)
ids = tok(text, add_special_tokens=False, return_tensors="pt").to(device)
with torch.no_grad():
logits = model(**ids, logits_to_keep=1).logits[0, -1]
letters = [tok(o["label"], add_special_tokens=False)["input_ids"][0] for o in opts]
return dict(zip(options, torch.softmax(logits[letters].float(), -1).tolist()))
state = "Blackjack. Player has 10 and 6 (hard 16). Dealer shows a 10. No surrender."
print(decide(state, "Best move for the player?", {"hit": "take another card", "stand": "keep the hand"}))
print(decide(state, "Will the player bust if they take one more card?", {"false": "No", "true": "Yes"})["true"])
For a yes/no question, use the options false: No and true: Yes in that order, as Ollama does. For a scale, use the level numbers 0, 1, 2, … as keys.
Held-out rows the model never saw in training, scored with the prompts Ollama sends. Chance is the score of picking an option at random. Brier and ECE measure how trustworthy the probabilities are; lower is better.
| set | rows | accuracy | chance | Brier | ECE |
|---|---|---|---|---|---|
| general decisions | 1500 | 68.1% | 35.7% | 0.127 | 0.032 |
| chess | 128 | 75.0% | 12.0% | 0.042 | 0.099 |
| poker | 127 | 84.3% | 42.5% | 0.087 | 0.046 |
| blackjack | 109 | 96.3% | 36.6% | 0.017 | 0.061 |
| grid paths | 101 | 93.1% | 55.5% | 0.068 | 0.212 |
| tic-tac-toe | 86 | 66.3% | 44.6% | 0.123 | 0.196 |
| Connect Four | 66 | 45.5% | 18.4% | 0.096 | 0.116 |
| Doom (rule-labelled) | 139 | 98.6% | 30.3% | 0.006 | 0.008 |
| FNAF (rule-labelled) | 76 | 100.0% | 19.3% | 0.000 | 0.004 |
| bets | 8 | 75.0% | 50.0% | 0.217 | 0.144 |
| all | 2340 | 73.9% | 34.7% | 0.100 | 0.033 |
The general decisions come from the validation split of tasksource-jev-typed-decisions: classification, NLI, QA, sentiment and similar tasks with human labels.
These are the full bf16 weights. The Q8_0 build on Ollama scores 72.6% overall on the same rows (ECE 0.042).
Speed, measured with a plain Transformers server in bf16 on an RTX 4070 Laptop GPU (median of 20 runs): 88 ms for one question, 116 ms for four, 447 ms for sixteen.
LoRA fine-tune of MiniCPM5-2B, merged into the weights afterwards. Rank 16 (alpha 32, dropout 0.05) on every linear layer, learning rate 1e-4 with a cosine schedule, batch 16, one epoch over about 55,000 examples, bf16, about five hours on an 8 GB RTX 4070 Laptop GPU. The loss is cross-entropy between the target distribution and the softmax over the allowed answer letters only; nothing else in the vocabulary is trained.
The data has two parts. 40,000 general decisions come from tasksource-jev-typed-decisions, with soft targets taken from annotator votes where the source has them. 15,360 game decisions were generated and labelled by program: chess positions scored by Stockfish at depth 13, poker spots by exact equity, blackjack hands by exact expected value (infinite deck, dealer stands on soft 17, double after split), simple bets by expected value, tic-tac-toe by minimax, Connect Four by win and block detection, grid movement by shortest path, and Doom and FNAF situations by hand-written rules. No game files were used.
After training, a temperature of 1.1 was fitted on the held-out rows to calibrate the probabilities, then folded into the output layer.
The Doom and FNAF labels come from hand-written rules, so the near-perfect scores mean the model learned those rules, not that it plays those games well. Connect Four and tic-tac-toe are its weakest games, and its confidence there runs higher than its accuracy. Bets has too few evaluation rows to say much.
This first version was trained on a plain-text question layout, where it scores 76.0% on the same rows. Ollama sends JSON, where it scores 73.9%; the table above uses Ollama’s format.
The general data comes from many sources with their own licenses, and each row of tasksource-jev-typed-decisions names its source and license. Some of those sources do not allow commercial use, so check them before using Aurelis commercially. The base model, MiniCPM5-2B, is Apache-2.0.