Store › model › Embeddings
WeMM-Embedding-2B-Quantized
by ewin-reg · source Hugging Face · updated 2026-09-24
other1.7 GB~3 GB RAMsource aliveunlabeled
WeMM-Embedding-2B-Quantized (Hybrid FP8 Attn/GDN + INT4-g16 MLP)
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:ewin-reg/WeMM-Embedding-2B-Quantized and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/ewin-reg/WeMM-Embedding-2B-Quantized
- License: other (custom license: read it at the source before installing)
- Requirements: about 3 GB of RAM, 1.7 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
sentence-transformerssafetensorsqwen3_5multimodalembeddingsretrievalfeature-extractionquantizationmixed-precisionw4a8fp8int4svdggufmrltext-embeddings
Numbers
- 1,927 downloads on Hugging Face
- 5 likes
- license other
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 21:00 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
model.safetensors | 1.7 GB |
From the source README
WeMM-Embedding-2B-Quantized (Hybrid FP8 Attn/GDN + INT4-g16 MLP)
Model Details
- Model Name: `WeMM-Embedding-2B-Quantized`
- Developer / Publisher: ewin-reg
- Base Architecture: `tencent/WeMM-Embedding-2B` (2.72B total parameters, Qwen3.5 hybrid architecture)
- Model Type: Omni-modal Foundation Embedding Model (Text, Image, Video)
- Quantization Scheme: Hybrid Curvature-Guided Mixed-Precision (Per-Token FP8 E4M3 Vocab + PAS-Guarded FP8 E4M3 Attention + Group-16 Symmetric INT4 MLPs)
- Format: Single Unified SafeTensors (`model.safetensors`, 1,791.14 MB / 1.749 GB)
- Embedding Dimensions: 2048 native (with Matryoshka Representation Learning down to 64 dims)
- Compatibility: 100% native Hugging Face and `SentenceTransformers` (`trust_remote_code=True`)
Intended Uses & Deployment Scope
Primary Use Cases
Card id model:hf:ewin-reg/WeMM-Embedding-2B-Quantized · collected 2026-10-02 21:00 UTC · JSON