LogiShell store Open app

Store › model › LLM

GLM-5.3-Flash-GGUF

by unsloth · source Hugging Face · updated 2026-09-06

mit9 MB~1 GB RAMsource aliveunlabeled

Unsloth Dynamic 3.0 achieves superior accuracy & outperforms other leading quants. To run, please use our llama.cpp PR or use the Unsloth Desktop app. You can now run GLM-5.3-Flash in our Unsloth Des…

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:unsloth/GLM-5.3-Flash-GGUF and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 20:52 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
BF16/GLM-5.3-Flash-BF16-00001-of-00014.ggufBF169 MB
BF16/GLM-5.3-Flash-BF16-00002-of-00014.ggufBF1645 GB
BF16/GLM-5.3-Flash-BF16-00003-of-00014.ggufBF1646 GB
BF16/GLM-5.3-Flash-BF16-00004-of-00014.ggufBF1646 GB
BF16/GLM-5.3-Flash-BF16-00005-of-00014.ggufBF1647 GB
BF16/GLM-5.3-Flash-BF16-00006-of-00014.ggufBF1646 GB
BF16/GLM-5.3-Flash-BF16-00007-of-00014.ggufBF1646 GB
BF16/GLM-5.3-Flash-BF16-00008-of-00014.ggufBF1646 GB
BF16/GLM-5.3-Flash-BF16-00009-of-00014.ggufBF1646 GB
BF16/GLM-5.3-Flash-BF16-00010-of-00014.ggufBF1646 GB
BF16/GLM-5.3-Flash-BF16-00011-of-00014.ggufBF1646 GB
BF16/GLM-5.3-Flash-BF16-00012-of-00014.ggufBF1646 GB
BF16/GLM-5.3-Flash-BF16-00013-of-00014.ggufBF1646 GB
BF16/GLM-5.3-Flash-BF16-00014-of-00014.ggufBF1646 GB
Q8_0/GLM-5.3-Flash-Q8_0-00001-of-00008.ggufQ8_09 MB
Q8_0/GLM-5.3-Flash-Q8_0-00002-of-00008.ggufQ8_046 GB
Q8_0/GLM-5.3-Flash-Q8_0-00003-of-00008.ggufQ8_046 GB
Q8_0/GLM-5.3-Flash-Q8_0-00004-of-00008.ggufQ8_046 GB
Q8_0/GLM-5.3-Flash-Q8_0-00005-of-00008.ggufQ8_047 GB
Q8_0/GLM-5.3-Flash-Q8_0-00006-of-00008.ggufQ8_046 GB
Q8_0/GLM-5.3-Flash-Q8_0-00007-of-00008.ggufQ8_046 GB
Q8_0/GLM-5.3-Flash-Q8_0-00008-of-00008.ggufQ8_039 GB
UD-IQ1_M/GLM-5.3-Flash-UD-IQ1_M-00001-of-00003.ggufIQ1_M9 MB
UD-IQ1_M/GLM-5.3-Flash-UD-IQ1_M-00002-of-00003.ggufIQ1_M47 GB
UD-IQ1_M/GLM-5.3-Flash-UD-IQ1_M-00003-of-00003.ggufIQ1_M44 GB
UD-IQ1_S/GLM-5.3-Flash-UD-IQ1_S-00001-of-00003.ggufIQ1_S9 MB
UD-IQ1_S/GLM-5.3-Flash-UD-IQ1_S-00002-of-00003.ggufIQ1_S46 GB
UD-IQ1_S/GLM-5.3-Flash-UD-IQ1_S-00003-of-00003.ggufIQ1_S40 GB
UD-IQ2_XXS/GLM-5.3-Flash-UD-IQ2_XXS-00001-of-00004.ggufIQ2_XXS9 MB
UD-IQ2_XXS/GLM-5.3-Flash-UD-IQ2_XXS-00002-of-00004.ggufIQ2_XXS46 GB
UD-IQ2_XXS/GLM-5.3-Flash-UD-IQ2_XXS-00003-of-00004.ggufIQ2_XXS46 GB
UD-IQ2_XXS/GLM-5.3-Flash-UD-IQ2_XXS-00004-of-00004.ggufIQ2_XXS2.5 GB
UD-IQ3_XXS/GLM-5.3-Flash-UD-IQ3_XXS-00001-of-00004.ggufIQ3_XXS9 MB
UD-IQ3_XXS/GLM-5.3-Flash-UD-IQ3_XXS-00002-of-00004.ggufIQ3_XXS46 GB
UD-IQ3_XXS/GLM-5.3-Flash-UD-IQ3_XXS-00003-of-00004.ggufIQ3_XXS46 GB
UD-IQ3_XXS/GLM-5.3-Flash-UD-IQ3_XXS-00004-of-00004.ggufIQ3_XXS21 GB
UD-IQ4_XS/GLM-5.3-Flash-UD-IQ4_XS-00001-of-00005.ggufIQ4_XS9 MB
UD-IQ4_XS/GLM-5.3-Flash-UD-IQ4_XS-00002-of-00005.ggufIQ4_XS47 GB
UD-IQ4_XS/GLM-5.3-Flash-UD-IQ4_XS-00003-of-00005.ggufIQ4_XS46 GB
UD-IQ4_XS/GLM-5.3-Flash-UD-IQ4_XS-00004-of-00005.ggufIQ4_XS46 GB
UD-IQ4_XS/GLM-5.3-Flash-UD-IQ4_XS-00005-of-00005.ggufIQ4_XS7.2 GB
UD-Q2_K_XL/GLM-5.3-Flash-UD-Q2_K_XL-00001-of-00004.ggufQ2_K_XL9 MB
UD-Q2_K_XL/GLM-5.3-Flash-UD-Q2_K_XL-00002-of-00004.ggufQ2_K_XL46 GB
UD-Q2_K_XL/GLM-5.3-Flash-UD-Q2_K_XL-00003-of-00004.ggufQ2_K_XL47 GB
UD-Q2_K_XL/GLM-5.3-Flash-UD-Q2_K_XL-00004-of-00004.ggufQ2_K_XL8.8 GB
UD-Q3_K_XL/GLM-5.3-Flash-UD-Q3_K_XL-00001-of-00004.ggufQ3_K_XL9 MB
UD-Q3_K_XL/GLM-5.3-Flash-UD-Q3_K_XL-00002-of-00004.ggufQ3_K_XL46 GB
UD-Q3_K_XL/GLM-5.3-Flash-UD-Q3_K_XL-00003-of-00004.ggufQ3_K_XL46 GB
UD-Q3_K_XL/GLM-5.3-Flash-UD-Q3_K_XL-00004-of-00004.ggufQ3_K_XL45 GB
UD-Q4_K_XL/GLM-5.3-Flash-UD-Q4_K_XL-00001-of-00006.ggufQ4_K_XL9 MB
UD-Q4_K_XL/GLM-5.3-Flash-UD-Q4_K_XL-00002-of-00006.ggufQ4_K_XL46 GB
UD-Q4_K_XL/GLM-5.3-Flash-UD-Q4_K_XL-00003-of-00006.ggufQ4_K_XL47 GB
UD-Q4_K_XL/GLM-5.3-Flash-UD-Q4_K_XL-00004-of-00006.ggufQ4_K_XL45 GB
UD-Q4_K_XL/GLM-5.3-Flash-UD-Q4_K_XL-00005-of-00006.ggufQ4_K_XL46 GB
UD-Q4_K_XL/GLM-5.3-Flash-UD-Q4_K_XL-00006-of-00006.ggufQ4_K_XL2.6 GB
UD-Q5_K_XL/GLM-5.3-Flash-UD-Q5_K_XL-00001-of-00006.ggufQ5_K_XL9 MB
UD-Q5_K_XL/GLM-5.3-Flash-UD-Q5_K_XL-00002-of-00006.ggufQ5_K_XL45 GB
UD-Q5_K_XL/GLM-5.3-Flash-UD-Q5_K_XL-00003-of-00006.ggufQ5_K_XL45 GB
UD-Q5_K_XL/GLM-5.3-Flash-UD-Q5_K_XL-00004-of-00006.ggufQ5_K_XL46 GB
UD-Q5_K_XL/GLM-5.3-Flash-UD-Q5_K_XL-00005-of-00006.ggufQ5_K_XL46 GB

From the source README

Read our How to Run GLM-5.3-Flash Guide!

Unsloth Dynamic 3.0 achieves superior accuracy & outperforms other leading quants.













To run, please use our llama.cpp PR or use the Unsloth Desktop app.
You can now run GLM-5.3-Flash in our Unsloth Desktop UI.
See below for 1-bit GLM-5.3-Flash (Low) run inside of Unsloth Desktop:

GLM-5.3-Flash

👋 Join our WeChat or Discord community.

📖 Check out the GLM-5.3-Flash blog and GLM-5 Technical report.

📍 Use GLM-5.3-Flash API services on Z.ai API Platform.

Introduction

We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute.

Footnotes

  • HLE w/ tools (full set): We use sampling parameters of `temperature=1.0` and `top_p=0.95` for evaluation, with a maximum generation length of `163,840` tokens. The evaluation is conducted with a maximum context length of `300,000` tokens, using a context management strategy. We use GPT-5.6-luna (medium) as the judge model.
  • **NL2…

Source: https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF

Card id model:hf:unsloth/GLM-5.3-Flash-GGUF · collected 2026-10-02 20:52 UTC · JSON