Store › model › Speech
Voxtral-Mini-4B-Realtime-2602-gguf
by handy-computer · source Hugging Face · updated 2026-09-15
apache-2.02.6 GB~4 GB RAMsource aliveunlabeled
Voxtral-Mini-4B-Realtime-2602: transcribe.cpp GGUF
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:handy-computer/Voxtral-Mini-4B-Realtime-2602-gguf and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/handy-computer/Voxtral-Mini-4B-Realtime-2602-gguf
- License: apache-2.0
- Requirements: about 4 GB of RAM, 2.6 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with whisper.cpp.
- Tags:
transcribe.cppggufasrspeech-to-textvoxtralaudio-llmstreamingmultilingualautomatic-speech-recognitionenfresderuzhja
Numbers
- 400,835 downloads on Hugging Face
- 6 likes
- license apache-2.0
- 2.6 GB for Voxtral-Mini-4B-Realtime-2602-Q4_K_M.gguf
Numbers as of 2026-10-02 20:59 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
Voxtral-Mini-4B-Realtime-2602-BF16.gguf | BF16 | 8.3 GB |
Voxtral-Mini-4B-Realtime-2602-F16.gguf | F16 | 8.3 GB |
Voxtral-Mini-4B-Realtime-2602-Q4_K_M.gguf | Q4_K_M | 2.6 GB |
Voxtral-Mini-4B-Realtime-2602-Q5_K_M.gguf | Q5_K_M | 3.1 GB |
Voxtral-Mini-4B-Realtime-2602-Q6_K.gguf | Q6_K | 3.4 GB |
Voxtral-Mini-4B-Realtime-2602-Q8_0.gguf | Q8_0 | 4.4 GB |
From the source README
Voxtral-Mini-4B-Realtime-2602: transcribe.cpp GGUF
GGUF conversions of mistralai/Voxtral-Mini-4B-Realtime-2602 for use
with transcribe.cpp.
Ported from upstream commit
2769294,
pinned 2026-06-06.
Validated against the Transformers reference at transcribe.cpp commit
483c122
on 2026-06-06.
Streaming audio-LLM speech-to-text. A ~970M causal audio encoder (left-pad
causal conv stem + 32-layer sliding-window RoPE transformer) feeds a
4-frame-group projector whose audio embeddings are added onto a ~3.4B
Ministral decoder (26 layers, GQA 32/8, NEOX RoPE) with delay-token latency
conditioning, emitting one text token per 80 ms audio slot (12.5 Hz). Takes a
16 kHz mono WAV and supports both incremental streaming (configurable
latency/quality via --stream-chunk-ms and --stream-voxtral-delay) and offline
transcription using an accuracy-first 2400 ms delay. Architecturally distinct
from the offline Voxtral 2507 family — own arch, streaming frontend, causal
encoder, additive audio fusion.
Downloads
| Quantization | Download | Size | WER (LibriSpeech test-clean) |
| --- | --- | ---: | ---: |
| BF16 | Voxtral-Mini-4B-Realtime-2602-BF16.gguf | 8.87 GB | 2.08% |
| F16 | Voxtral-Mini-4B-Realtime-2602-F16.gguf | 8.88 GB | 2.09% |
| Q8_0 | [Voxtral-Mini-4B-Realtime-2602-Q8_0.gguf](https://huggingface.co/handy-computer/Voxtral-Mini-4B-Realtime-2602-gguf/resolve/main/Voxtral-Mini-4B-Re…
Source: https://huggingface.co/handy-computer/Voxtral-Mini-4B-Realtime-2602-gguf
Card id model:hf:handy-computer/Voxtral-Mini-4B-Realtime-2602-gguf · collected 2026-10-02 20:59 UTC · JSON