15 2 weeks ago

The Ministral 3 family is designed for edge deployment, capable of running on a wide range of hardware. This version is the converted model from Huggingface with full FB16 precision.

ollama run beeble/ministral-3-3b-instruct-2512-fb16

Details

2 weeks ago

d6408f4a6ba6 · 6.9GB

mistral3
·
3.43B
·
BF16
{ "num_ctx": 16384, "temperature": 0.15 }

Readme

Ministral 3 3B Instruct 2512 — FB16 Build

Full-precision FB16 (16-bit float) build of Mistral AI’s Ministral 3 3B Instruct 2512 — no quantization, weights stored exactly as released. Output is (nearly) identical in quality to model version hosted by Mistral itself.

FB16 vs. Standard Q4_0

The default Ollama tag is usually a 4-bit build (q4_0). Key differences:

FB16 (this build) Q4_0 (default)
Quality Bit-exact, zero quantization error ~4.25 bits/weight; on a small 3B model the quality loss is proportionally larger (weaker instruction adherence, more reasoning slips, worse calibration)
Size ~6–7 GB ~2–2.5 GB
Speed Slower (weights are 4× larger to stream) Typically 2–3× more tokens/sec

When to use which: Q4_0 for edge hardware, limited RAM/VRAM, and fast casual chat. FB16 when fidelity matters — benchmarking, structured output (JSON, code, tool calls), agentic pipelines where small errors cascade, or generating reference data.

Disclaimer

  • This build was independently created. The model provider is not acting on behalf of Mistral AI and has no affiliation with Mistral AI whatsoever.
  • Model weights are subject to Mistral AI’s license — review it before commercial use.

License

Model weights: Mistral AI. This build and README are provided as-is, without warranty.