{"v":1,"id":"model:hf:handy-computer/Voxtral-Small-24B-2507-gguf","slug":"model-handy-computer-voxtral-small-24b-2507-gguf","kind":"model","category":"speech","title":"Voxtral-Small-24B-2507-gguf","summary":"Voxtral-Small-24B-2507: transcribe.cpp GGUF","source":{"provider":"hf","ref":"handy-computer/Voxtral-Small-24B-2507-gguf","url":"https://huggingface.co/handy-computer/Voxtral-Small-24B-2507-gguf","rev":"4ab5a0708a16619b2b5b98f14b20333e0de17943","fetchedAt":"2026-10-02T20:59:45.411Z","etag":"W/\"d65-c6YX+oPt3kpNNSrhS71STy4c1eg\""},"author":{"name":"handy-computer","url":"https://huggingface.co/handy-computer"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":94503,"likes":0,"takenAt":"2026-10-02T20:59:45.411Z"},"tags":["transcribe.cpp","gguf","asr","speech-to-text","voxtral","audio-llm","multilingual","automatic-speech-recognition","en","fr","de","es","it","pt","nl","hi"],"pipeline":"automatic-speech-recognition","links":{"github":"mistralai/mistral-inference"},"updatedAt":"2026-09-15T07:05:45.000Z","collectedAt":"2026-10-02T20:59:45.411Z","review":{"numbers":["94,503 downloads on Hugging Face","0 likes","license apache-2.0","13 GB for Voxtral-Small-24B-2507-Q4_K_M.gguf"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:45.411Z","http":200},"description":"# Voxtral-Small-24B-2507: transcribe.cpp GGUF\n\nGGUF conversions of [mistralai/Voxtral-Small-24B-2507](https://huggingface.co/mistralai/Voxtral-Small-24B-2507) for use\nwith [transcribe.cpp](https://github.com/handy-computer/transcribe.cpp).\n\nPorted from upstream commit\n[da5b424](https://huggingface.co/mistralai/Voxtral-Small-24B-2507/commit/da5b424),\npinned 2026-06-05.\nValidated against the Transformers reference at transcribe.cpp commit\n[dac22fa](https://github.com/handy-computer/transcribe.cpp/tree/dac22fa)\non 2026-06-05.\n\nOffline audio-LLM speech-to-text and speech translation. A Whisper-large-v3\nbidirectional audio encoder feeds a 4-frame-group projector (375 audio tokens\nper 30 s chunk) into a Mistral-Small-24B causal LM (40 layers, GQA 32/8, NEOX\nRoPE, SwiGLU) via audio-token injection. Takes a 16 kHz mono WAV and produces a\ntranscript via greedy decoding. The larger sibling of Voxtral Mini 3B — same\nencoder, projector, frontend, and tokenizer, with a scaled-up decoder.\n\n## Downloads\n\n| Quantization | Download | Size | WER (LibriSpeech test-clean) |\n| --- | --- | ---: | ---: |\n| BF16 | [Voxtral-Small-24B-2507-BF16.gguf](https://huggingface.co/handy-computer/Voxtral-Small-24B-2507-gguf/resolve/main/Voxtral-Small-24B-2507-BF16.gguf) | 48.54 GB | 1.56% |\n| F16 | [Voxtral-Small-24B-2507-F16.gguf](https://huggingface.co/handy-computer/Voxtral-Small-24B-2507-gguf/resolve/main/Voxtral-Small-24B-2507-F16.gguf) | 48.55 GB | 1.57% |\n| Q8_0 | [Voxtral-Small-24B-2507-Q8_0.gguf](https://huggingface.co/handy-computer/Voxtral-Small-24B-2507-gguf/resolve/main/Voxtral-Small-24B-2507-Q8_0.gguf) | 25.81 GB | 1.56% |\n| Q6_K | [Voxtral-Small-24B-2507-Q6_K.gguf](https://huggingface.co/handy-computer/Voxtral-Small-24B-2507-gguf/resolve/main/Voxtral-Small-24B-2507-Q6_K.gguf) | 19.94 GB | 1.58% |\n| Q5_K_M | [Voxtral-Small-24B-2507-Q5_K_M.gguf](https://huggingface.co/handy-computer/Voxtral-Small-24B-2507-gguf/re…\n\nSource: https://huggingface.co/handy-computer/Voxtral-Small-24B-2507-gguf","install":{"kind":"model","hfId":"handy-computer/Voxtral-Small-24B-2507-gguf","gated":false,"format":"gguf","files":[{"name":"Voxtral-Small-24B-2507-BF16.gguf","size":48537285088,"quant":"BF16","sha256":"5c38ff2d6dbb413d70dec37430825ff129862536f74149e015b3bc42c11c5745"},{"name":"Voxtral-Small-24B-2507-F16.gguf","size":48548098528,"quant":"F16","sha256":"d5cc73553d9422dada176dd02c1920dc72a888a01eed34e20816b0bab8db207e"},{"name":"Voxtral-Small-24B-2507-Q4_K_M.gguf","size":14302261728,"quant":"Q4_K_M","sha256":"0da1b866c0eb2ce805678c18a7fb1e413320e4ad3ccbef7160707b65bae937a3"},{"name":"Voxtral-Small-24B-2507-Q5_K_M.gguf","size":17138659808,"quant":"Q5_K_M","sha256":"a53f73a5f63b7663fe155977616a636acac567dac898f6ef82ca519017e8b6e7"},{"name":"Voxtral-Small-24B-2507-Q6_K.gguf","size":19936473568,"quant":"Q6_K","sha256":"e85950b9576d3913d2664abaac1cdcb3a967dbbc8a60febacaf137a7f4d216f0"},{"name":"Voxtral-Small-24B-2507-Q8_0.gguf","size":25810383328,"quant":"Q8_0","sha256":"ab0964350131990a364dca53e3cac5f4d9cc176dce2f273066be3ef71e252fb2"}],"totalBytes":174273162048,"suggestedFile":"Voxtral-Small-24B-2507-Q4_K_M.gguf","requirements":{"ramGb":16,"diskBytes":14302261728,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["whisper.cpp"],"command":"lsh models install hf:handy-computer/Voxtral-Small-24B-2507-gguf"}}