Store › model › Speech
canary-qwen-2.5b-gguf
by handy-computer · source Hugging Face · updated 2026-09-15
cc-by-4.01.6 GB~3 GB RAMsource aliveunlabeled
GGUF conversions of nvidia/canary-qwen-2.5b for use with transcribe.cpp.
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:handy-computer/canary-qwen-2.5b-gguf and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/handy-computer/canary-qwen-2.5b-gguf
- License: cc-by-4.0
- Requirements: about 3 GB of RAM, 1.6 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with whisper.cpp.
- Tags:
transcribe.cppggufasrspeech-to-textcanarycanary-qwensalmaudio-llmfastconformerqwen3automatic-speech-recognitionenconversational
Numbers
- 55,992 downloads on Hugging Face
- 1 likes
- license cc-by-4.0
- 1.6 GB for canary-qwen-2.5b-Q4_K_M.gguf
Numbers as of 2026-10-02 21:00 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
canary-qwen-2.5b-BF16.gguf | BF16 | 4.7 GB |
canary-qwen-2.5b-F16.gguf | F16 | 4.7 GB |
canary-qwen-2.5b-Q4_K_M.gguf | Q4_K_M | 1.6 GB |
canary-qwen-2.5b-Q5_K_M.gguf | Q5_K_M | 1.8 GB |
canary-qwen-2.5b-Q6_K.gguf | Q6_K | 2.1 GB |
canary-qwen-2.5b-Q8_0.gguf | Q8_0 | 2.6 GB |
From the source README
canary-qwen-2.5b: transcribe.cpp GGUF
GGUF conversions of nvidia/canary-qwen-2.5b for use
with transcribe.cpp.
Ported from upstream commit
b1469e1,
pinned 2026-05-15.
Validated against the NeMo SALM 2.7.3 reference at transcribe.cpp commit
6f6c699
on 2026-05-16.
Offline English speech-to-text. NeMo SALM (Speech-Augmented Language
Model): a FastConformer audio encoder (32 layers, `d_model=1024`) feeds
audio embeddings into a Qwen3-1.7B causal LM (28 layers,
`hidden_size=2048`) via audio-token injection at a sentinel position in
the prompt. English only. Takes a 16 kHz mono WAV and produces a
transcript via greedy decoding.
Downloads
| Quantization | Download | Size | WER (LibriSpeech test-clean) |
| --- | --- | ---: | ---: |
| BF16 | canary-qwen-2.5b-BF16.gguf | 5.08 GB | 1.63% |
| F16 | canary-qwen-2.5b-F16.gguf | 5.08 GB | 1.63% |
| Q8_0 | canary-qwen-2.5b-Q8_0.gguf | 2.80 GB | 1.63% |
| Q6_K | canary-qwen-2.5b-Q6_K.gguf | 2.21 GB | 1.63% |
| Q5_K_M | canary-qwen-2.5b-Q5_K_M.gguf | 1.98 GB | 1.63% |
| Q4_K_M | canary-qwen-2.5b-Q4_K_M.gguf | 1.74 GB |…
Source: https://huggingface.co/handy-computer/canary-qwen-2.5b-gguf
Card id model:hf:handy-computer/canary-qwen-2.5b-gguf · collected 2026-10-02 21:00 UTC · JSON