4 Downloads Updated 7 hours ago
ollama pull benjaminj/octen-embedding-8b:q8_0
Octen-Embedding-8B is a text embedding model developed by Octen for semantic search and retrieval. It is fine-tuned (LoRA) from Qwen/Qwen3-Embedding-8B and supports 100+ languages and code.
Unofficial GGUF build of Octen/Octen-Embedding-8B (revision
5adcfa29) for Ollama. All credit for the model goes to the Octen team.
| Model | Dim | Mean (Public) | Mean (Private) | Mean (Task) |
|---|---|---|---|---|
| Octen-Embedding-8B | 4096 | 0.7953 | 0.8157 | 0.8045 |
| voyage-3-large | 1024 | 0.7434 | 0.8277 | 0.7812 |
| gemini-embedding-001 | 3072 | 0.7218 | 0.8075 | 0.7602 |
| Qwen3-Embedding-8B | 4096 | 0.7310 | 0.7838 | 0.7547 |
| text-embedding-3-large | 3072 | 0.6110 | 0.7130 | 0.6567 |
| bge-m3 | 1024 | 0.5216 | 0.6726 | 0.5893 |
| Tag | Quantization | Size | Cosine vs. BF16 reference (min / mean) |
|---|---|---|---|
latest, q4_K_M |
Q4_K_M | 4.7 GB | 0.987 / 0.989 |
q8_0 |
Q8_0 | 8.0 GB | 0.998 / 0.9998 |
Checked against the original BF16 weights run with sentence-transformers 5.7. Retrieval ranking and token counts are identical.
ollama pull benjaminj/octen-embedding-8b
Ollama does not apply the sentence-transformers prompts: prefix them yourself.
Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery: + your query
(literal newline before Query:). You can adapt the instruction to your task."- " (dash + space) + your text. Upstream recommends this to avoid an
upstream Qwen3-Embedding issue with unprefixed documents.curl http://localhost:11434/api/embed -d '{
"model": "benjaminj/octen-embedding-8b",
"input": [
"Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery:What is the capital of Australia?",
"- Canberra is the capital city of Australia."
]
}'
Semantic search and retrieval, document similarity and clustering, question answering, cross-lingual retrieval, classification with embeddings.
Converted with llama.cpp convert_hf_to_gguf.py (commit b9acf138) to F16, then quantized with llama-quantize.
GGUF metadata matches the official qwen3-embedding:8b: pooling_type=3 (last token), rope.freq_base=1e6, add_eos_token=true.
@misc{octen2025rteb,
title={Octen Series: Optimizing Embedding Models to #1 on RTEB Leaderboard},
author={Octen Team},
year={2025},
url={https://octen-team.github.io/octen_blog/posts/octen-rteb-first-place/}
}