{"v":1,"id":"model:hf:ewin-reg/WeMM-Embedding-2B-Quantized","slug":"model-ewin-reg-wemm-embedding-2b-quantized","kind":"model","category":"embedding","title":"WeMM-Embedding-2B-Quantized","summary":"WeMM-Embedding-2B-Quantized (Hybrid FP8 Attn/GDN + INT4-g16 MLP)","source":{"provider":"hf","ref":"ewin-reg/WeMM-Embedding-2B-Quantized","url":"https://huggingface.co/ewin-reg/WeMM-Embedding-2B-Quantized","rev":"524bbe61e1b8ce4673e4e17dcf6a8694a3c45993","fetchedAt":"2026-10-02T21:00:20.572Z","etag":"W/\"3355-iXx0Ko03+sh3WgL6QKNgVgf71Yk\""},"author":{"name":"ewin-reg","url":"https://huggingface.co/ewin-reg"},"license":{"spdx":null,"raw":"other","open":null,"note":"custom license: read it at the source before installing"},"metrics":{"downloads":1927,"downloadsWeek":327838,"likes":5,"takenAt":"2026-10-02T21:00:20.572Z"},"tags":["sentence-transformers","safetensors","qwen3_5","multimodal","embeddings","retrieval","feature-extraction","quantization","mixed-precision","w4a8","fp8","int4","svd","gguf","mrl","text-embeddings","image-embedding","video-embedding","cross-modal","custom_code","en","zh","multilingual","model-index"],"pipeline":"feature-extraction","links":{"github":"ggml-org/llama.cpp","npm":"node-llama-cpp"},"updatedAt":"2026-09-24T17:10:27.000Z","collectedAt":"2026-10-02T21:00:20.572Z","review":{"numbers":["1,927 downloads on Hugging Face","5 likes","license other","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T21:00:20.572Z","http":200},"description":"# WeMM-Embedding-2B-Quantized (Hybrid FP8 Attn/GDN + INT4-g16 MLP)\n\n[](https://huggingface.co/ewin-reg/WeMM-Embedding-2B-Quantized)\n\n[](https://huggingface.co/tencent/WeMM-Embedding-2B)\n\n[](https://huggingface.co/ewin-reg/WeMM-Embedding-2B-Quantized)\n\n[](https://sbert.net/)\n\n## Model Details\n\n- **Model Name**: `WeMM-Embedding-2B-Quantized`\n\n- **Developer / Publisher**: ewin-reg\n\n- **Base Architecture**: [`tencent/WeMM-Embedding-2B`](https://huggingface.co/tencent/WeMM-Embedding-2B) (2.72B total parameters, Qwen3.5 hybrid architecture)\n\n- **Model Type**: Omni-modal Foundation Embedding Model (Text, Image, Video)\n\n- **Quantization Scheme**: Hybrid Curvature-Guided Mixed-Precision (Per-Token FP8 E4M3 Vocab + PAS-Guarded FP8 E4M3 Attention + Group-16 Symmetric INT4 MLPs)\n\n- **Format**: Single Unified SafeTensors (`model.safetensors`, 1,791.14 MB / 1.749 GB)\n\n- **Embedding Dimensions**: 2048 native (with Matryoshka Representation Learning down to 64 dims)\n\n- **Compatibility**: 100% native Hugging Face and `SentenceTransformers` (`trust_remote_code=True`)\n\n---\n\n## Intended Uses & Deployment Scope\n\n### Primary Use Cases\n\n- **High-Throughput Multimodal Retrieval**: Semantic document search, zero-shot text-to-image ranking, and video clip retrieval.\n\n- **Edge & Constrained Deployments**: Production vector databases and edge servers constrained to 1.5 GB – 2.0 GB memory budgets.\n\n- **Native Python Pipelines**: Pure Python execution via `SentenceTransformer(\"ewin-reg/WeMM-Embedding-2B-Quantized\", trust_remote_code=True)` without external C++ runtimes or specialized GGUF fork dependencies.\n\n- **Flexible Vector Indexing (MRL)**: Dynamic dimension truncation (from 2048 down to 1024, 512, 256, 128, or 64 dimensions) for extreme vector indexing efficiency.\n\n### Out-of-Scope & Limitations\n\n- **Generative Text Output**: The causal language modeling head has been replaced with mean-pooled embedding projections; it d…\n\nSource: https://huggingface.co/ewin-reg/WeMM-Embedding-2B-Quantized","install":{"kind":"model","hfId":"ewin-reg/WeMM-Embedding-2B-Quantized","gated":false,"format":"safetensors","files":[{"name":"model.safetensors","size":1878150064,"sha256":"f8e8ce9332ca2bbe7225f789b1aa54289fec6d4205e99ea3e6f4692c519f41d5"}],"totalBytes":1878150064,"requirements":{"ramGb":3,"diskBytes":1878150064,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:ewin-reg/WeMM-Embedding-2B-Quantized"}}