smtek/ BigBang-v1:map-k4v

227 2 weeks ago

A 35B-A3B agentic model based on Qwen3.6-35B-A3B, designed for long‑horizon search, software engineering, scientific research, and AI research + map-k4v

tools
ollama run smtek/BigBang-v1:map-k4v

Details

2 weeks ago

f5925f5d8b83 · 22GB

qwen35moe
·
35.5B
·
Q4_K_M
{{- range $i, $_ := .Messages }}{{- $last := eq (len (slice $.Messages $i)) 1 }}{{- if eq .Role "use
{ "draft_ngram_map_k4v_min_hits": 1, "draft_ngram_map_k4v_size_m": 48, "draft_ngram_map_

Readme

BigBang-v1

GGUF quantizations of endless-frontier/BigBang-v1 for Ollama, with context windows tuned per VRAM budget.

Model Description

BigBang-v1 is a 35B‑A3B Mixture‑of‑Experts model built on Qwen3.6‑35B‑A3B. It has 35B total parameters but only 3B activated per inference, striking a great balance between capability and efficiency.

In 8 representative benchmarks covering long‑horizon search, software engineering, scientific reasoning, and AI research, BigBang‑v1 achieved the highest average score among selected 35B‑class models. It even outperformed DeepSeek V4 Pro Preview (1.6T) on FrontierScience Research, Humanity’s Last Exam, PaperBench(Code‑Dev), and BioMysteryBench‑HD.

Feature Details
Base model endless‑frontier/BigBang‑v1 (Qwen3.6‑35B‑A3B)
GGUF source bartowski/endless‑frontier_BigBang‑v1‑GGUF
Architecture 35B‑A3B hybrid MoE, 40 layers, 10 full-attention layers
Hardware Runs comfortably on 24GB VRAM cards (e.g., RTX 3090⁄4090)

Models

Tag Quant Size Context VRAM target Use case
latest 4-bit 21.9 GB 160K 24 GB Default
Q4_K_S 4-bit 21.1 GB 96K 24 GB Recommended (100% GPU, 1 GB headroom)
IQ4_XS 4-bit 19.3 GB 256K 24 GB Smaller & faster, worse tool calling
Q2_K-16gb 2-bit 13.1 GB 112K 16 GB Fits 16 GB VRAM

Tool-calling note (from our benchmark): all tags handle flat-schema tools (weather, calculator, search, parallel calls) and the full call→answer loop. The one gap is nested-object tool schemas (e.g. book_flight with a nested passenger object): Q2_K-16gb crashes Ollama with an HTTP 500 on those, while latest, Q4_K_S, and IQ4_XS handle them. If you rely on nested tool args, use a 4-bit tag or flatten the schema.

Recommended Optimization

By default, Ollama allocates the KV-Cache in f16, which will exceed 24GB VRAM over long contexts. To run stably, set these environment variables in your system (/etc/systemd/system/ollama.service.d/override.conf on Linux):

[Service]
Environment="OLLAMA_FLASH_ATTENTION=1"
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"
  • Flash Attention: Dramatically reduces context memory footprint.
  • q8_0 Cache: Compresses context tensors without the severe logic degradation caused by 4-bit (q4_0) cache quantization.

After editing, reload and restart:

sudo systemctl daemon-reload
sudo systemctl restart ollama

Important: custom Ollama build

The map-k4v, MTP3 + map-k4v, and DFlash2 + map-k4v tags use speculative-decoding fields that stock Ollama does not forward to llama.cpp. Stock Ollama still works for each clean baseline tag, but it does not activate these custom draft routes.

Build and run the custom runtime from the Ollama source. Build notes, local patches, and updates are maintained on Samuel Ishida’s GitHub.

cmake -B build .
cmake --build build --parallel 8
./ollama serve

For this tag set, use the custom launcher from the patched checkout:

scripts/ollama-ngram.sh build --full
OLLAMA_NGRAM_GPU_MASK=1 scripts/ollama-ngram.sh start-custom

The published tags remain usable as normal Ollama models. The custom build is needed only when users want the configured speculative route. Hardware, driver, GPU, cache type, and context size affect results; numbers below are local measurements, not universal guarantees.

Configuration:

draft_spec_type=ngram-map-k4v
draft_num_predict=0
draft_ngram_map_k4v_size_n=12
draft_ngram_map_k4v_size_m=48
draft_ngram_map_k4v_min_hits=1

map-k4v is the validated BigBang winner. It does not add a second token stream; it proposes candidates that the target verifies.

Short agentic result

Two agentic coding turns, 128 output tokens per turn, 50K-token synthetic repository, 65,536-token context, q4_0 K/V cache, batch 256, greedy seed 42, one cold run per arm, full Vulkan offload on RX 7900 XTX device 1:

Arm Aggregate decode tok/s Speedup Parity
Baseline 93.41 1.00x PASS
map-k4v n=12 102.41 1.10x PASS
map-k4v n=24 94.06 1.01x PASS
map-k4v n=32 89.78 0.96x PASS

Use n=12, m=48, min_hits=1. It was fastest and retained exact output parity against baseline.

The compatible BigBang DFlash2 test tag is not included in this push. Its Qwen3.5-compatible draft loaded and activated, but measured 56.10–60.62 tok/s and failed strict parity. The original Qwen3.8 DFlash2 head was incompatible with BigBang’s qwen35moe architecture.

Links