{"v":1,"id":"model:hf:ewin-reg/WeMM-Embedding-2B-GGUF","slug":"model-ewin-reg-wemm-embedding-2b-gguf","kind":"model","category":"embedding","title":"WeMM-Embedding-2B-GGUF","summary":"Model Card for WeMM-Embedding-2B-GGUF (Unofficial)","source":{"provider":"hf","ref":"ewin-reg/WeMM-Embedding-2B-GGUF","url":"https://huggingface.co/ewin-reg/WeMM-Embedding-2B-GGUF","rev":"a219b89a6b677e114c125696a2f5e4da95098fd6","fetchedAt":"2026-10-02T20:59:55.539Z","etag":"W/\"15ff-MsxQBDXCQgGWkTgIOljA6+quyI8\""},"author":{"name":"ewin-reg","url":"https://huggingface.co/ewin-reg"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":2739,"downloadsWeek":327838,"likes":2,"takenAt":"2026-10-02T20:59:55.539Z"},"tags":["gguf","qwen3_5","llama-cpp","unsloth","sentence-transformers","text-embeddings","multimodal-embedding","feature-extraction","code-search","flatquant","schurscale","quantized","q4_k_m","q5_k_m","q6_k","q8_0","int4","int8","custom_code","en","zh","model-index","endpoints_compatible"],"pipeline":"feature-extraction","links":{"github":"ggml-org/llama.cpp","npm":"node-llama-cpp"},"updatedAt":"2026-09-08T00:24:20.000Z","collectedAt":"2026-10-02T20:59:55.539Z","review":{"numbers":["2,739 downloads on Hugging Face","2 likes","license apache-2.0","1.5 GB for gguf/wemm-embedding-2b-q4_k_m.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:55.539Z","http":200},"description":"# Model Card for WeMM-Embedding-2B-GGUF (Unofficial)\n\n> [!IMPORTANT]\n> **New Native SafeTensors Release Available:**\n> For pure Python and native SentenceTransformers execution with full multimodal support (Text, Vision ViT, and Video), a smaller footprint (**1.437 GB** vs 1.453 GB Q4_K_M), and preserved attention softmax without requiring llama.cpp forks, use the new native release:\n>\n> 👉 [**ewin-reg/WeMM-Embedding-2B-Quantized**](https://huggingface.co/ewin-reg/WeMM-Embedding-2B-Quantized)\n\n[](https://huggingface.co/ewin-reg/WeMM-Embedding-2B-Quantized)\n\nCommunity (Unofficial) GGUF, PyTorch INT8, and 2026 Research INT4 checkpoints for [**tencent/WeMM-Embedding-2B**](https://huggingface.co/tencent/WeMM-Embedding-2B), an efficient 2B-parameter hybrid architecture (18 Mamba SSM linear-attention layers + 6 full-attention layers) engineered for high-throughput code retrieval and multimodal embedding.\n\nCompatible with **`llama.cpp`**, **`Ollama`**, **`LM Studio`**, **`Unsloth`**, **`vLLM`**, and **`sentence-transformers`**.\n\n---\n\n## Model Details\n\n### Model Description\n\n- **Developed by:** Tencent (Base Model); Quantized by [ewinregirgojr](https://huggingface.co/ewinregirgojr)\n- **Model Type:** Hybrid Linear-Attention Mamba SSM + Softmax Attention Embedding Model\n- **Language(s) (NLP):** English, Chinese, Multilingual, and Programming Languages (Python, TypeScript, JavaScript, C++, Rust, Go, SQL, Shell)\n- **License:** Apache-2.0\n- **Base Model:** [`tencent/WeMM-Embedding-2B`](https://huggingface.co/tencent/WeMM-Embedding-2B)\n- **Embedding Dimension:** 2,048 dimensions (L2-normalized dense float vector)\n\n### Model Sources\n\n- **Base Model Repository:** [tencent/WeMM-Embedding-2B](https://huggingface.co/tencent/WeMM-Embedding-2B)\n- **Quantized Repository:** [ewinregirgojr/WeMM-Embedding-2B-GGUF](https://huggingface.co/ewinregirgojr/WeMM-Embedding-2B-GGUF)\n- **Inference Framework:** [llama.cpp](https://github…\n\nSource: https://huggingface.co/ewin-reg/WeMM-Embedding-2B-GGUF","install":{"kind":"model","hfId":"ewin-reg/WeMM-Embedding-2B-GGUF","gated":false,"format":"gguf","files":[{"name":"gguf/wemm-embedding-2b-q4_k_m.gguf","size":1559762176,"quant":"Q4_K_M","sha256":"6900f992cafed2a32fcc8bc53dd95e5b6ac55af72bec7c4b901fc9ef9946989c"},{"name":"gguf/wemm-embedding-2b-q5_k_m.gguf","size":1759994624,"quant":"Q5_K_M","sha256":"3808320ade23486d83ecc32a7b0587500f1c73dd72b3458c2d51c45f393c6cdb"},{"name":"gguf/wemm-embedding-2b-q6_k.gguf","size":1972741600,"quant":"Q6_K","sha256":"c17640dc239055e891d68990a8bf3632958b32cb0baca4821b38e9676409c968"},{"name":"gguf/wemm-embedding-2b-q8_0.gguf","size":2551289888,"quant":"Q8_0","sha256":"d7d264dd4488105232668e59a7ea30bc5d86b78a0539309e77d217f292c5b09d"},{"name":"pytorch/model_int8.safetensors","size":3233535108,"sha256":"c95b57e187f1486e8588e73e074d51c4c49bcd0ca04d8b115d765096ea66f6bf"},{"name":"research/model_research_int4.safetensors","size":3369579584,"sha256":"3ffa26c14035a1929c09cff27caf9cf14c44241a95e47013107749d0197a08e0"},{"name":"research/model_schurscale_int4.safetensors","size":3369579584,"sha256":"c6f6a60c289d2e7001b1ebebbb74d5ead66462e27326d9f57f87481922c65c60"},{"name":"research/wemm-embedding-2b-research-q4.gguf","size":1559762176,"sha256":"fc1e853db7bc72c0e57e58d232c25d1636612b77dc5cd1dd0f9f6b1b0069716e"},{"name":"research/wemm-embedding-2b-schurscale-q4.gguf","size":1559762176,"sha256":"92cbc772295988f98158f57bdf3bcefd06ccb9564204b57d9079b8c9c6a8a012"}],"totalBytes":20936006916,"suggestedFile":"gguf/wemm-embedding-2b-q4_k_m.gguf","requirements":{"ramGb":3,"diskBytes":1559762176,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:ewin-reg/WeMM-Embedding-2B-GGUF"}}