Store › model › Speech
Fun-ASR-MLT-Nano-2512-gguf
by handy-computer · source Hugging Face · updated 2026-09-15
other531 MB~2 GB RAMsource aliveunlabeled
Fun-ASR-MLT-Nano-2512: transcribe.cpp GGUF
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:handy-computer/Fun-ASR-MLT-Nano-2512-gguf and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/handy-computer/Fun-ASR-MLT-Nano-2512-gguf
- License: other (custom license: read it at the source before installing) · text
- Requirements: about 2 GB of RAM, 531 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with whisper.cpp.
- Tags:
transcribe.cppggufasrspeech-to-textfun-asr-nanofun-asr-mlt-nanosense-voice-encoderqwen3audio-llmmultilingualautomatic-speech-recognitionzhenyuejako
Numbers
- 33,103 downloads on Hugging Face
- 1 likes
- license other
- 0.5 GB for Fun-ASR-MLT-Nano-2512-Q4_K_M.gguf
- 15,195 npm downloads a week for nodejs-whisper
- latest nodejs-whisper@0.3.1
Numbers as of 2026-10-02 21:00 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
Fun-ASR-MLT-Nano-2512-BF16.gguf | BF16 | 1.6 GB |
Fun-ASR-MLT-Nano-2512-F16.gguf | F16 | 1.6 GB |
Fun-ASR-MLT-Nano-2512-Q4_K_M.gguf | Q4_K_M | 531 MB |
Fun-ASR-MLT-Nano-2512-Q5_K_M.gguf | Q5_K_M | 602 MB |
Fun-ASR-MLT-Nano-2512-Q6_K.gguf | Q6_K | 659 MB |
Fun-ASR-MLT-Nano-2512-Q8_0.gguf | Q8_0 | 850 MB |
From the source README
Fun-ASR-MLT-Nano-2512: transcribe.cpp GGUF
GGUF conversions of FunAudioLLM/Fun-ASR-MLT-Nano-2512 for use
with transcribe.cpp.
Ported from upstream commit
cf67a93,
pinned 2026-05-06.
Validated against the FunASR reference at transcribe.cpp commit
f094d28
on 2026-05-06.
Offline speech-to-text covering 31 languages, with focused optimization
on East and Southeast Asian languages: Chinese, English, Cantonese,
Japanese, Korean, Vietnamese, Indonesian, Thai, Malay, Filipino, plus
Arabic, Hindi, and 19 European languages (Bulgarian, Croatian, Czech,
Danish, Dutch, Estonian, Finnish, Greek, Hungarian, Irish, Latvian,
Lithuanian, Maltese, Polish, Portuguese, Romanian, Slovak, Slovenian,
Swedish). Same architecture as Fun-ASR-Nano-2512 (~800M trainable
parameters: frozen SenseVoiceEncoderSmall + 2-layer audio adaptor +
bundled Qwen3-0.6B LLM); trained on a smaller multilingual corpus
("hundreds of thousands of hours" per the model card, vs Nano's
"tens of millions"). Takes a 16 kHz mono WAV and emits text. Not
streaming, no translation, no timestamps. ITN (inverse text
normalization) is supported by the model and exposed via the
`--itn` CLI flag and `transcribe_funasr_nano_params { use_itn }`
in the library API.
Downloads
| Quantization | Download | Size | WER (LibriSpeech test-clean) |
| --- | --- | ---: | ---: |
| BF16 | Fun-ASR-MLT-Nano-2512-BF16.gguf | 1.67 GB | 1.74% |
| F16 | Fun-ASR-MLT-Nano-2512-F16.gguf | 1.67 GB |…
Source: https://huggingface.co/handy-computer/Fun-ASR-MLT-Nano-2512-gguf
Card id model:hf:handy-computer/Fun-ASR-MLT-Nano-2512-gguf · collected 2026-10-02 21:00 UTC · JSON