maternion/ spark-x2.5:4b-q4_K_M

156 3 days ago

Spark-X2.5-4B is an efficient, open-source AI model built for chat, writing, coding, reasoning, tools, and agents, supporting 1M-token context and 200+ languages.

4b
ollama run maternion/spark-x2.5:4b-q4_K_M

Details

3 days ago

a7e7cc3781a6 · 2.6GB

spark2_5
·
4.11B
·
Q4_K_M
{ "temperature": 1, "top_k": 0, "top_p": 0.95 }

Readme

Spark-X2.5

xhtoken.png

Introduction

We are introducing Spark-X2.5-4B and Spark-X2.5-1.7B, two compact, general-purpose language models designed to make capable AI more practical, efficient, and accessible. The models deliver strong performance across a broad range of everyday tasks—including conversation, writing, translation, reasoning, coding, tool use, and agentic workflows—achieving leading results among open-source models of comparable size. Spark-X2.5 combines an efficiency-oriented architecture with native context windows of up to 1M tokens, and support for more than 200 languages.

Technical Highlights

  • Efficient Architecture and Native 1M-token Context: The models use a hybrid attention architecture that combines one full-attention layer with three sliding-window attention layers. This design substantially reduces the computational overhead typically associated with long-context models while natively supporting a context window of up to 1M tokens.

  • Strong Coding and Agent Capabilities: The models are deeply integrated with popular agent harnesses, including Codex, Claude Code, OpenClaw, and Hermes. They deliver state-of-the-art performance among models of comparable size across everyday coding, agentic workflows, reasoning, and instruction-following tasks.

  • Broad Hardware and Software Compatibility: The models support a wide range of hardware platforms, including NVIDIA, Huawei, Hygon, and HOUMO.AI. They are compatible with leading inference frameworks such as vLLM, SGLang, llama.cpp, and MLX, and can be deployed through Ollama and LM Studio.

  • Advanced Training Algorithms: The models were trained on Huawei Ascend clusters. Large-scale reinforcement learning and post-training techniques such as MOPD significantly enhance their reasoning, coding, agentic, and instruction-following capabilities.

model-benchmark-comparison.svg

Model Overview

For agent tasks, balancing performance, inference speed, and cache usage has long been a key bottleneck limiting model performance. Spark-X2.5 systematically integrates and optimizes mature attention technologies, combining sliding-window attention (SWA) with a hybrid full-attention architecture. This approach leverages the strengths of both mechanisms while avoiding the limitations of relying on a single structure, achieving an effective balance among performance, inference efficiency, and KV-cache size—thereby improving its practicality and effectiveness across real-world deployment scenarios.

spark25-hybrid-architecture-light.png

Training Methods

Spark-X2.5 is pretrained on approximately 20 trillion tokens from a diverse corpus spanning web pages, books, academic publications, code, and encyclopedic materials. Particular attention is paid to data quality, domain coverage, and the sampling weights assigned to different data categories. Extensive data-mixture studies are conducted to determine an effective balance among mathematics, logic, code, and other high-value domains. This enables the models to acquire broad general knowledge while developing stronger capabilities in complex reasoning and code generation. Long-context capability is developed through a dedicated training stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens.

Post-training begins with supervised fine-tuning on a carefully curated corpus. This stage establishes robust instruction following, structured generation, and task-completion, while providing a stable policy initialization for reinforcement learning. We subsequently apply large-scale reinforcement learning across several capability domains, including language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following. This process yields a set of domain-specialized teacher policies, whose complementary strengths are consolidated into a single deployable model through MOPD.

post_training_pipeline.svg

Benchmarks

We evaluate our models and compare them with leading on-device models of similar size across a broad range of tasks, including agent, code, math, general and knowledge.

Benchmark Spark‑X2.5‑4B Spark‑X2.5‑1.7B Qwen3.5‑9B Qwen3.5‑4B Qwen3.5‑2B Gemma4‑12B Gemma4‑E4B Gemma4‑E2B
Agent
BFCL‑V4 65.1 46.9 66.1* 50.3* 43.6* 37.4 36.9 30.2
τ²‑bench 75.1 65.3 79.1* 79.9* 48.8* 69.0* 42.2* 24.5*
τ³‑bench 30.4 20.1 9.3 6.7 4.1 13.3 10.1 8.8
MCP‑Atlas 54.6 23.4 47.4* 40.8* 14.8 30.5* 15.0* 12.6
MCP‑Mark 14.2 2.3 13.4 12.5
Workspace Bench 31.2 18.9 25.5 21.3 7.7
VitaBench2.0 25.2 8.3 15.6 18.2 5.2 12.4 4.8 4.4
BrowseComp 40.9 29.7 8.3 14.3 3.1 10.0 8.3 3.7
Code
SWE‑Bench Pro 44.4 10.4 33.8* 29.4* 1.9 21.9* 4.0*
SWE‑Bench Verified 41.6 28.3 53.1* 38.8* 6.8 44.2* 14.0*
SWE‑Bench Multilingual 53.3 23.3 43.3 27.7 5.0 32.5*
SciCode 34.7 18.2 32.7* 24.0 6.0 39.8 27.5 20.5
Math
Gaokao 2026 133.4 114.8 135.5 130.3 94.0 130.6 102.4 81.8
AIME 2026 90.7 69.4 88.2 83.0 30.8 82.1* 42.5* 37.5*
HMMT Feb 2026 81.2 48.4 70.8 69.7 21.5 65.6 34.2 20.5
IMO‑AnswerBench 74.2 45.4 69.8 68.5 57.2 26.9 22.6
General & Knowledge
IFEval 93.0 89.5 91.5* 89.8* 78.6* 94.8 45.3 34.8
IFBench 75.0 66.3 64.5 59.2 41.3* 73.5* 44.0* 22.7
AA‑LCR 56.3 24.3 63.0* 57.0* 25.6* 55.3* 34.7 18.3
HLE 12.3 6.3 14.3 8.6 2.1 13.1 3.9 2.5
GPQA 67.4 43.8 77.2 67.2 44.6 72.8 54.5 43.8
  • * denotes reported results from publicly‑released model cards / papers and - denotes scores not yet available.
  • All evaluations are conducted in thinking mode. The recommended sampling parameters for Spark-X2.5 are temperature=1.0, top_p=0.95, and top_k=-1.
  • Gaokao 2026 consists of the five 2026 Chinese GAOKAO examinations (National I, National II, Beijing, Shanghai, Tianjin), each graded out of 150 points.