Store › model › Speech
Voxtral-Small-24B-2507-gguf
by handy-computer · source Hugging Face · updated 2026-09-15
apache-2.013 GB~16 GB RAMsource aliveunlabeled
Voxtral-Small-24B-2507: transcribe.cpp GGUF
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:handy-computer/Voxtral-Small-24B-2507-gguf and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/handy-computer/Voxtral-Small-24B-2507-gguf
- License: apache-2.0
- Requirements: about 16 GB of RAM, 13 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with whisper.cpp.
- Tags:
transcribe.cppggufasrspeech-to-textvoxtralaudio-llmmultilingualautomatic-speech-recognitionenfrdeesitptnlhi
Numbers
- 94,503 downloads on Hugging Face
- 0 likes
- license apache-2.0
- 13 GB for Voxtral-Small-24B-2507-Q4_K_M.gguf
Numbers as of 2026-10-02 20:59 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
Voxtral-Small-24B-2507-BF16.gguf | BF16 | 45 GB |
Voxtral-Small-24B-2507-F16.gguf | F16 | 45 GB |
Voxtral-Small-24B-2507-Q4_K_M.gguf | Q4_K_M | 13 GB |
Voxtral-Small-24B-2507-Q5_K_M.gguf | Q5_K_M | 16 GB |
Voxtral-Small-24B-2507-Q6_K.gguf | Q6_K | 19 GB |
Voxtral-Small-24B-2507-Q8_0.gguf | Q8_0 | 24 GB |
From the source README
Voxtral-Small-24B-2507: transcribe.cpp GGUF
GGUF conversions of mistralai/Voxtral-Small-24B-2507 for use
with transcribe.cpp.
Ported from upstream commit
da5b424,
pinned 2026-06-05.
Validated against the Transformers reference at transcribe.cpp commit
dac22fa
on 2026-06-05.
Offline audio-LLM speech-to-text and speech translation. A Whisper-large-v3
bidirectional audio encoder feeds a 4-frame-group projector (375 audio tokens
per 30 s chunk) into a Mistral-Small-24B causal LM (40 layers, GQA 32/8, NEOX
RoPE, SwiGLU) via audio-token injection. Takes a 16 kHz mono WAV and produces a
transcript via greedy decoding. The larger sibling of Voxtral Mini 3B — same
encoder, projector, frontend, and tokenizer, with a scaled-up decoder.
Downloads
| Quantization | Download | Size | WER (LibriSpeech test-clean) |
| --- | --- | ---: | ---: |
| BF16 | Voxtral-Small-24B-2507-BF16.gguf | 48.54 GB | 1.56% |
| F16 | Voxtral-Small-24B-2507-F16.gguf | 48.55 GB | 1.57% |
| Q8_0 | Voxtral-Small-24B-2507-Q8_0.gguf | 25.81 GB | 1.56% |
| Q6_K | Voxtral-Small-24B-2507-Q6_K.gguf | 19.94 GB | 1.58% |
| Q5_K_M | [Voxtral-Small-24B-2507-Q5_K_M.gguf](https://huggingface.co/handy-computer/Voxtral-Small-24B-2507-gguf/re…
Source: https://huggingface.co/handy-computer/Voxtral-Small-24B-2507-gguf
Card id model:hf:handy-computer/Voxtral-Small-24B-2507-gguf · collected 2026-10-02 20:59 UTC · JSON