LogiShell store Open app

Store › model › Speech

Fun-ASR-MLT-Nano-2512-gguf

by handy-computer · source Hugging Face · updated 2026-09-15

other531 MB~2 GB RAMsource aliveunlabeled

Fun-ASR-MLT-Nano-2512: transcribe.cpp GGUF

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:handy-computer/Fun-ASR-MLT-Nano-2512-gguf and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 21:00 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
Fun-ASR-MLT-Nano-2512-BF16.ggufBF161.6 GB
Fun-ASR-MLT-Nano-2512-F16.ggufF161.6 GB
Fun-ASR-MLT-Nano-2512-Q4_K_M.ggufQ4_K_M531 MB
Fun-ASR-MLT-Nano-2512-Q5_K_M.ggufQ5_K_M602 MB
Fun-ASR-MLT-Nano-2512-Q6_K.ggufQ6_K659 MB
Fun-ASR-MLT-Nano-2512-Q8_0.ggufQ8_0850 MB

From the source README

Fun-ASR-MLT-Nano-2512: transcribe.cpp GGUF

GGUF conversions of FunAudioLLM/Fun-ASR-MLT-Nano-2512 for use
with transcribe.cpp.

Ported from upstream commit
cf67a93,
pinned 2026-05-06.
Validated against the FunASR reference at transcribe.cpp commit
f094d28
on 2026-05-06.

Offline speech-to-text covering 31 languages, with focused optimization
on East and Southeast Asian languages: Chinese, English, Cantonese,
Japanese, Korean, Vietnamese, Indonesian, Thai, Malay, Filipino, plus
Arabic, Hindi, and 19 European languages (Bulgarian, Croatian, Czech,
Danish, Dutch, Estonian, Finnish, Greek, Hungarian, Irish, Latvian,
Lithuanian, Maltese, Polish, Portuguese, Romanian, Slovak, Slovenian,
Swedish). Same architecture as Fun-ASR-Nano-2512 (~800M trainable
parameters: frozen SenseVoiceEncoderSmall + 2-layer audio adaptor +
bundled Qwen3-0.6B LLM); trained on a smaller multilingual corpus
("hundreds of thousands of hours" per the model card, vs Nano's
"tens of millions"). Takes a 16 kHz mono WAV and emits text. Not
streaming, no translation, no timestamps. ITN (inverse text
normalization) is supported by the model and exposed via the
`--itn` CLI flag and `transcribe_funasr_nano_params { use_itn }`
in the library API.

Downloads

| Quantization | Download | Size | WER (LibriSpeech test-clean) |
| --- | --- | ---: | ---: |
| BF16 | Fun-ASR-MLT-Nano-2512-BF16.gguf | 1.67 GB | 1.74% |
| F16 | Fun-ASR-MLT-Nano-2512-F16.gguf | 1.67 GB |…

Source: https://huggingface.co/handy-computer/Fun-ASR-MLT-Nano-2512-gguf

Card id model:hf:handy-computer/Fun-ASR-MLT-Nano-2512-gguf · collected 2026-10-02 21:00 UTC · JSON