Store › model › Embeddings
WeMM-Embedding-2B-GGUF
by ewin-reg · source Hugging Face · updated 2026-09-08
apache-2.01.5 GB~3 GB RAMsource aliveunlabeled
Model Card for WeMM-Embedding-2B-GGUF (Unofficial)
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:ewin-reg/WeMM-Embedding-2B-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/ewin-reg/WeMM-Embedding-2B-GGUF
- License: apache-2.0
- Requirements: about 3 GB of RAM, 1.5 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
ggufqwen3_5llama-cppunslothsentence-transformerstext-embeddingsmultimodal-embeddingfeature-extractioncode-searchflatquantschurscalequantizedq4_k_mq5_k_mq6_kq8_0
Numbers
- 2,739 downloads on Hugging Face
- 2 likes
- license apache-2.0
- 1.5 GB for gguf/wemm-embedding-2b-q4_k_m.gguf
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:59 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
gguf/wemm-embedding-2b-q4_k_m.gguf | Q4_K_M | 1.5 GB |
gguf/wemm-embedding-2b-q5_k_m.gguf | Q5_K_M | 1.6 GB |
gguf/wemm-embedding-2b-q6_k.gguf | Q6_K | 1.8 GB |
gguf/wemm-embedding-2b-q8_0.gguf | Q8_0 | 2.4 GB |
pytorch/model_int8.safetensors | 3.0 GB | |
research/model_research_int4.safetensors | 3.1 GB | |
research/model_schurscale_int4.safetensors | 3.1 GB | |
research/wemm-embedding-2b-research-q4.gguf | 1.5 GB | |
research/wemm-embedding-2b-schurscale-q4.gguf | 1.5 GB |
From the source README
Model Card for WeMM-Embedding-2B-GGUF (Unofficial)
> [!IMPORTANT]
> New Native SafeTensors Release Available:
> For pure Python and native SentenceTransformers execution with full multimodal support (Text, Vision ViT, and Video), a smaller footprint (1.437 GB vs 1.453 GB Q4_K_M), and preserved attention softmax without requiring llama.cpp forks, use the new native release:
>
> 👉 ewin-reg/WeMM-Embedding-2B-Quantized
Community (Unofficial) GGUF, PyTorch INT8, and 2026 Research INT4 checkpoints for tencent/WeMM-Embedding-2B, an efficient 2B-parameter hybrid architecture (18 Mamba SSM linear-attention layers + 6 full-attention layers) engineered for high-throughput code retrieval and multimodal embedding.
Compatible with `llama.cpp`, `Ollama`, `LM Studio`, `Unsloth`, `vLLM`, and `sentence-transformers`.
Model Details
Model Description
- Developed by: Tencent (Base Model); Quantized by ewinregirgojr
- Model Type: Hybrid Linear-Attention Mamba SSM + Softmax Attention Embedding Model
- Language(s) (NLP): English, Chinese, Multilingual, and Programming Languages (Python, TypeScript, JavaScript, C++, Rust, Go, SQL, Shell)
- License: Apache-2.0
- Base Model: `tencent/WeMM-Embedding-2B`
- Embedding Dimension: 2,048 dimensions (L2-normalized dense float vector)
Model Sources
- Base Model Repository: tencent/WeMM-Embedding-2B
- Quantized Repository: ewinregirgojr/WeMM-Embedding-2B-GGUF
- Inference Framework: [llama.cpp](https://github…
Source: https://huggingface.co/ewin-reg/WeMM-Embedding-2B-GGUF
Card id model:hf:ewin-reg/WeMM-Embedding-2B-GGUF · collected 2026-10-02 20:59 UTC · JSON