15 Downloads Updated 2 weeks ago
ollama run beeble/ministral-3-3b-instruct-2512-fb16
Full-precision FB16 (16-bit float) build of Mistral AI’s Ministral 3 3B Instruct 2512 — no quantization, weights stored exactly as released. Output is (nearly) identical in quality to model version hosted by Mistral itself.
The default Ollama tag is usually a 4-bit build (q4_0). Key differences:
| FB16 (this build) | Q4_0 (default) | |
|---|---|---|
| Quality | Bit-exact, zero quantization error | ~4.25 bits/weight; on a small 3B model the quality loss is proportionally larger (weaker instruction adherence, more reasoning slips, worse calibration) |
| Size | ~6–7 GB | ~2–2.5 GB |
| Speed | Slower (weights are 4× larger to stream) | Typically 2–3× more tokens/sec |
When to use which: Q4_0 for edge hardware, limited RAM/VRAM, and fast casual chat. FB16 when fidelity matters — benchmarking, structured output (JSON, code, tool calls), agentic pipelines where small errors cascade, or generating reference data.
Model weights: Mistral AI. This build and README are provided as-is, without warranty.