127 4 days ago

whisper.cpp large-v3 turbo · 99 languages · translation to English

ollama run mannix/whisper:large-v3-turbo

Models

View all →

Readme

Whisper large-v3 turbo — transcription and translation, for xOllama

whisper.cpp large-v3 turbo (Q8_0) · 99 languages · translation to English · OpenAI /v1/audio/transcriptions and /v1/audio/translations

Runs on xOllama, not on stock ollama. Stock ollama pulls this model but cannot show or run it: it carries only media engines. xOllama serves it through an opencoti engine with media support (dev build 2610022033001/b105 or later, set with XOLLAMA_ENGINE_PATH). The engine pinned by the current xOllama release has no media engines and refuses the model at load, naming what it lacks.

Speech to text in 99 languages, and translation of any of them to English, from OpenAI’s turbo variant of Whisper large-v3.

Use

xollama pull mannix/whisper:large-v3-turbo
curl http://localhost:22434/v1/audio/transcriptions \
  -F model=mannix/whisper:large-v3-turbo -F file=@meeting.mp3
curl http://localhost:22434/v1/audio/translations \
  -F model=mannix/whisper:large-v3-turbo -F file=@intervista.m4a

wav, mp3 and m4a/aac are read as they are. The reply is OpenAI’s {"text": …} (the template’s default response_format is json).

Measured

a short clip, RTX 3090 (b97) about 1 s, exact
eight TTS clips of one sentence, RTX PRO 6000 (b112) all transcribed word for word

Tags

Tag Weights Size
large-v3-turbo ggml-large-v3-turbo-q8_0.bin 874 MB

Source

ggerganov/whisper.cpp · License: MIT. Docs: xOllama media.