4 7 hours ago

Octen-Embedding-8B, #1 on the RTEB leaderboard. 4096-dim multilingual embeddings fine-tuned from Qwen3-Embedding-8B for retrieval (legal, finance, healthcare, code).

embedding
ollama pull benjaminj/octen-embedding-8b:q8_0

Details

7 hours ago

fd0f412f5746 · 8.0GB

qwen3
·
7.57B
·
Q8_0
Apache License Version 2.0, January 2004 http://www.apache.org/licenses/ TERMS AND CONDITIONS FOR US
{{ .Prompt }}

Readme

Octen-Embedding-8B

Octen-Embedding-8B is a text embedding model developed by Octen for semantic search and retrieval. It is fine-tuned (LoRA) from Qwen/Qwen3-Embedding-8B and supports 100+ languages and code.

Unofficial GGUF build of Octen/Octen-Embedding-8B (revision 5adcfa29) for Ollama. All credit for the model goes to the Octen team.

Key highlights

  • #1 on the RTEB leaderboard (as of January 12, 2026): Mean (Task) 0.8045, Public 0.7953, Private 0.8157
  • Vertical domains: legal, finance, healthcare, code
  • Long context: up to 32,768 tokens (40,960 max sequence length)
  • Multilingual: 100+ languages, cross-lingual and code retrieval

RTEB leaderboard (excerpt, from the upstream model card)

Model Dim Mean (Public) Mean (Private) Mean (Task)
Octen-Embedding-8B 4096 0.7953 0.8157 0.8045
voyage-3-large 1024 0.7434 0.8277 0.7812
gemini-embedding-001 3072 0.7218 0.8075 0.7602
Qwen3-Embedding-8B 4096 0.7310 0.7838 0.7547
text-embedding-3-large 3072 0.6110 0.7130 0.6567
bge-m3 1024 0.5216 0.6726 0.5893

Model details

  • Base model: Qwen3-Embedding-8B (7.6B parameters, 36 layers)
  • Embedding dimension: 4096, last-token pooling, L2-normalized
  • Max sequence length: 40,960 tokens
  • License: Apache 2.0 (same as upstream)

Tags

Tag Quantization Size Cosine vs. BF16 reference (min / mean)
latest, q4_K_M Q4_K_M 4.7 GB 0.987 / 0.989
q8_0 Q8_0 8.0 GB 0.998 / 0.9998

Checked against the original BF16 weights run with sentence-transformers 5.7. Retrieval ranking and token counts are identical.

Usage

ollama pull benjaminj/octen-embedding-8b

Ollama does not apply the sentence-transformers prompts: prefix them yourself.

  • Query: Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery: + your query (literal newline before Query:). You can adapt the instruction to your task.
  • Document: "- " (dash + space) + your text. Upstream recommends this to avoid an upstream Qwen3-Embedding issue with unprefixed documents.
curl http://localhost:11434/api/embed -d '{
  "model": "benjaminj/octen-embedding-8b",
  "input": [
    "Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery:What is the capital of Australia?",
    "- Canberra is the capital city of Australia."
  ]
}'

Recommended use cases

Semantic search and retrieval, document similarity and clustering, question answering, cross-lingual retrieval, classification with embeddings.

Limitations

  • Performance may vary across domains and languages
  • Documents longer than 40K tokens must be truncated
  • Built for retrieval, not for text generation

Build

Converted with llama.cpp convert_hf_to_gguf.py (commit b9acf138) to F16, then quantized with llama-quantize. GGUF metadata matches the official qwen3-embedding:8b: pooling_type=3 (last token), rope.freq_base=1e6, add_eos_token=true.

Citation

@misc{octen2025rteb,
  title={Octen Series: Optimizing Embedding Models to #1 on RTEB Leaderboard},
  author={Octen Team},
  year={2025},
  url={https://octen-team.github.io/octen_blog/posts/octen-rteb-first-place/}
}