Store › model › Speech
VibeVoice-ASR-BitNet
by microsoft · source Hugging Face · updated 2026-07-24
mit947 MB~2 GB RAMsource aliveunlabeled
VibeVoice-ASR-BitNet is a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs — no GPU required. Through heterogeneous quantization, the model is compressed from 4.62 G…
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:microsoft/VibeVoice-ASR-BitNet and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
- License: mit
- Requirements: about 2 GB of RAM, 947 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with whisper.cpp.
- Tags:
ggmlsafetensorsggufvibevoiceASRquantizationcpu-inferencebitnetmultilingualautomatic-speech-recognitionenzhfritkopt
Numbers
- 61,271 downloads on Hugging Face
- 202 likes
- license mit
- 0.9 GB for vibeasr-lm-i2_s-embed-q6_k.gguf
- 15,195 npm downloads a week for nodejs-whisper
- latest nodejs-whisper@0.3.1
Numbers as of 2026-10-02 21:00 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
model-00001-of-00003.safetensors | 4.7 GB | |
model-00002-of-00003.safetensors | 4.6 GB | |
model-00003-of-00003.safetensors | 1.2 GB | |
vibeasr-lm-i2_s-embed-q6_k.gguf | Q6_K | 947 MB |
vibeasr-vae-encoder-i8_s.gguf | 671 MB |
From the source README
VibeVoice-ASR-BitNet
VibeVoice-ASR-BitNet is a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs — no GPU required. Through heterogeneous quantization, the model is compressed from 4.62 GB to 1.58 GB while achieving 1.6–2.3× faster inference than Whisper.cpp with real-time capability (RTF
➡️ Report: VibeVoice-ASR-BitNet Technical Report
➡️ Base Model: microsoft/VibeVoice-ASR
🔥 Key Features
- ⚡ Real-time on CPU — RTF
| Component | FP16 | Quantized | Method | Compression |
|:-:|:-:|:-:|:-:|:-:|
| VAE Tokenizer | 1.31 GB | 0.65 GB | I8\_S | 2.0× |
| LM Decoder | 3.32 GB | 0.92 GB | I2\_S + Q6\_K | 3.6× |
| Total | 4.62 GB | 1.58 GB | — | 2.9× |
Evaluation
Inference Speed
| Threads | 1 | 2 | 3 | 4 | 6 | 8 |
|:-:|:-:|:-:|:-:|:-:|:-:|:-:|
| RTF | 1.98 | 1.08 | 0.77 | 0.63 | 0.49 | 0.42 |
| vs. Whisper.cpp | 2.28× | 2.12× | 1.86× | 1.86× | 1.71× | 1.55× |
> Benchmarked on AMD EPYC 7V13 (AVX2+FMA) with 20s audio. Bold = RTF
| Benchmark | VibeVoice-ASR-7B | VibeVoice-ASR-BitNet | Parakeet | Whisper | SenseVoice | FunASR |
|:-:|:-:|:-:|:-:|:-:|:-:|:-:|
| MLC-EN | 7.82 | 8.25 | 8.40 | 13.57 | 12.39 | 11.36 |
| MLC-FR | 16.03 | 17.41 | — | — | — | — |
| MLC-IT | 15.67 | 17.23 | — | — | — | — |
| MLC-KO | 9.83 | 11.15 | — | — | — | — |
| MLC-PT | 22.41 | 24.87 | — | — | — | — |
| MLC-VI | 20.15 | 22.38 | — | — | — | — |
| AISHELL4 | 19.83 | 27.45 | — | — | 22.52 | 20.41 |
| AMI-ihm | 17.42 | 21.36 | 21.92 | 27.07 | 30.81 | 32.07 |
| AMI-sdm | 24.18 | 25.87 | 26.33 | 36.92 | 48.11 | 40.17 |
| AliMeeting | 36.21 | 40.58…
Source: https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
Card id model:hf:microsoft/VibeVoice-ASR-BitNet · collected 2026-10-02 21:00 UTC · JSON