Store › model › Speech
Qwen3-ASR-1.7B-gguf
by handy-computer · source Hugging Face · updated 2026-09-15
apache-2.01.2 GB~2 GB RAMsource aliveunlabeled
GGUF conversions of Qwen/Qwen3-ASR-1.7B for use with transcribe.cpp.
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:handy-computer/Qwen3-ASR-1.7B-gguf and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/handy-computer/Qwen3-ASR-1.7B-gguf
- License: apache-2.0
- Requirements: about 2 GB of RAM, 1.2 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with whisper.cpp.
- Tags:
transcribe.cppggufasrspeech-to-textqwen3qwen3-asrmultilingualautomatic-speech-recognitionzhenyueardefrespt
Numbers
- 195,152 downloads on Hugging Face
- 3 likes
- license apache-2.0
- 1.2 GB for Qwen3-ASR-1.7B-Q4_K_M.gguf
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:59 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
Qwen3-ASR-1.7B-BF16.gguf | BF16 | 3.8 GB |
Qwen3-ASR-1.7B-F16.gguf | F16 | 3.8 GB |
Qwen3-ASR-1.7B-Q4_K_M.gguf | Q4_K_M | 1.2 GB |
Qwen3-ASR-1.7B-Q5_K_M.gguf | Q5_K_M | 1.4 GB |
Qwen3-ASR-1.7B-Q6_K.gguf | Q6_K | 1.6 GB |
Qwen3-ASR-1.7B-Q8_0.gguf | Q8_0 | 2.0 GB |
From the source README
Qwen3-ASR-1.7B: transcribe.cpp GGUF
GGUF conversions of Qwen/Qwen3-ASR-1.7B for use
with transcribe.cpp.
Ported from upstream commit
7278e1e,
pinned 2026-04-19.
Validated against the qwen_asr 0.0.6 reference at transcribe.cpp commit
3f61df7
on 2026-04-20.
Offline multilingual speech-to-text. Same audio-LLM architecture as the
0.6B variant (bidirectional audio encoder feeding a Qwen3 causal LM with
audio-token injection), wider: encoder `d_model=1024` (16 heads), LM
`hidden_size=2048`, `intermediate_size=6144`. Auto-detects the audio's
language across 30 languages and emits the transcript in that language.
Takes a 16 kHz mono WAV; explicit language hints are not supported at
this time.
Downloads
| Quantization | Download | Size | WER (LibriSpeech test-clean) |
| --- | --- | ---: | ---: |
| BF16 | Qwen3-ASR-1.7B-BF16.gguf | 4.08 GB | 1.62% |
| F16 | Qwen3-ASR-1.7B-F16.gguf | 4.09 GB | 1.62% |
| Q8_0 | Qwen3-ASR-1.7B-Q8_0.gguf | 2.19 GB | 1.62% |
| Q6_K | Qwen3-ASR-1.7B-Q6_K.gguf | 1.69 GB | 1.65% |
| Q5_K_M | Qwen3-ASR-1.7B-Q5_K_M.gguf | 1.52 GB | 1.65% |
| Q4_K_M | [Qwen3-ASR-1.7B-Q4_K_M.gguf](https://huggingface.co/handy-computer/Qwen3-ASR-1.7B-gguf/resolve/main/Qwen3-ASR-1.7B-Q4_K_M…
Source: https://huggingface.co/handy-computer/Qwen3-ASR-1.7B-gguf
Card id model:hf:handy-computer/Qwen3-ASR-1.7B-gguf · collected 2026-10-02 20:59 UTC · JSON