227 Downloads Updated 2 weeks ago
ollama run smtek/BigBang-v1
Updated 1 month ago
1 month ago
01b7f2026e1f · 22GB
GGUF quantizations of endless-frontier/BigBang-v1 for Ollama, with context windows tuned per VRAM budget.
BigBang-v1 is a 35B‑A3B Mixture‑of‑Experts model built on Qwen3.6‑35B‑A3B. It has 35B total parameters but only 3B activated per inference, striking a great balance between capability and efficiency.
In 8 representative benchmarks covering long‑horizon search, software engineering, scientific reasoning, and AI research, BigBang‑v1 achieved the highest average score among selected 35B‑class models. It even outperformed DeepSeek V4 Pro Preview (1.6T) on FrontierScience Research, Humanity’s Last Exam, PaperBench(Code‑Dev), and BioMysteryBench‑HD.
| Feature | Details |
|---|---|
| Base model | endless‑frontier/BigBang‑v1 (Qwen3.6‑35B‑A3B) |
| GGUF source | bartowski/endless‑frontier_BigBang‑v1‑GGUF |
| Architecture | 35B‑A3B hybrid MoE, 40 layers, 10 full-attention layers |
| Hardware | Runs comfortably on 24GB VRAM cards (e.g., RTX 3090⁄4090) |
| Tag | Quant | Size | Context | VRAM target | Use case |
|---|---|---|---|---|---|
latest |
4-bit | 21.9 GB | 160K | 24 GB | Default |
Q4_K_S |
4-bit | 21.1 GB | 96K | 24 GB | Recommended (100% GPU, 1 GB headroom) |
IQ4_XS |
4-bit | 19.3 GB | 256K | 24 GB | Smaller & faster, worse tool calling |
Q2_K-16gb |
2-bit | 13.1 GB | 112K | 16 GB | Fits 16 GB VRAM |
Tool-calling note (from our benchmark): all tags handle flat-schema tools (weather, calculator, search, parallel calls) and the full call→answer loop. The one gap is nested-object tool schemas (e.g.
book_flightwith a nestedpassengerobject):Q2_K-16gbcrashes Ollama with an HTTP 500 on those, whilelatest,Q4_K_S, andIQ4_XShandle them. If you rely on nested tool args, use a 4-bit tag or flatten the schema.
By default, Ollama allocates the KV-Cache in f16, which will exceed 24GB VRAM over long contexts. To run stably, set these environment variables in your system (/etc/systemd/system/ollama.service.d/override.conf on Linux):
[Service]
Environment="OLLAMA_FLASH_ATTENTION=1"
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"
q4_0) cache quantization.After editing, reload and restart:
sudo systemctl daemon-reload
sudo systemctl restart ollama
The map-k4v, MTP3 + map-k4v, and DFlash2 + map-k4v tags use speculative-decoding fields that stock Ollama does not forward to llama.cpp. Stock Ollama still works for each clean baseline tag, but it does not activate these custom draft routes.
Build and run the custom runtime from the Ollama source. Build notes, local patches, and updates are maintained on Samuel Ishida’s GitHub.
cmake -B build .
cmake --build build --parallel 8
./ollama serve
For this tag set, use the custom launcher from the patched checkout:
scripts/ollama-ngram.sh build --full
OLLAMA_NGRAM_GPU_MASK=1 scripts/ollama-ngram.sh start-custom
The published tags remain usable as normal Ollama models. The custom build is needed only when users want the configured speculative route. Hardware, driver, GPU, cache type, and context size affect results; numbers below are local measurements, not universal guarantees.
Configuration:
draft_spec_type=ngram-map-k4v
draft_num_predict=0
draft_ngram_map_k4v_size_n=12
draft_ngram_map_k4v_size_m=48
draft_ngram_map_k4v_min_hits=1
map-k4v is the validated BigBang winner. It does not add a second token
stream; it proposes candidates that the target verifies.
Two agentic coding turns, 128 output tokens per turn, 50K-token synthetic repository, 65,536-token context, q4_0 K/V cache, batch 256, greedy seed 42, one cold run per arm, full Vulkan offload on RX 7900 XTX device 1:
| Arm | Aggregate decode tok/s | Speedup | Parity |
|---|---|---|---|
| Baseline | 93.41 | 1.00x | PASS |
map-k4v n=12 |
102.41 | 1.10x | PASS |
map-k4v n=24 |
94.06 | 1.01x | PASS |
map-k4v n=32 |
89.78 | 0.96x | PASS |
Use n=12, m=48, min_hits=1. It was fastest and retained exact output
parity against baseline.
The compatible BigBang DFlash2 test tag is not included in this push. Its
Qwen3.5-compatible draft loaded and activated, but measured 56.10–60.62 tok/s
and failed strict parity. The original Qwen3.8 DFlash2 head was incompatible
with BigBang’s qwen35moe architecture.