LogiShell store Open app

Store › model › Embeddings

WeMM-Embedding-2B-GGUF

by ewin-reg · source Hugging Face · updated 2026-09-08

apache-2.01.5 GB~3 GB RAMsource aliveunlabeled

Model Card for WeMM-Embedding-2B-GGUF (Unofficial)

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:ewin-reg/WeMM-Embedding-2B-GGUF and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 20:59 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
gguf/wemm-embedding-2b-q4_k_m.ggufQ4_K_M1.5 GB
gguf/wemm-embedding-2b-q5_k_m.ggufQ5_K_M1.6 GB
gguf/wemm-embedding-2b-q6_k.ggufQ6_K1.8 GB
gguf/wemm-embedding-2b-q8_0.ggufQ8_02.4 GB
pytorch/model_int8.safetensors3.0 GB
research/model_research_int4.safetensors3.1 GB
research/model_schurscale_int4.safetensors3.1 GB
research/wemm-embedding-2b-research-q4.gguf1.5 GB
research/wemm-embedding-2b-schurscale-q4.gguf1.5 GB

From the source README

Model Card for WeMM-Embedding-2B-GGUF (Unofficial)

> [!IMPORTANT]
> New Native SafeTensors Release Available:
> For pure Python and native SentenceTransformers execution with full multimodal support (Text, Vision ViT, and Video), a smaller footprint (1.437 GB vs 1.453 GB Q4_K_M), and preserved attention softmax without requiring llama.cpp forks, use the new native release:
>
> 👉 ewin-reg/WeMM-Embedding-2B-Quantized

Community (Unofficial) GGUF, PyTorch INT8, and 2026 Research INT4 checkpoints for tencent/WeMM-Embedding-2B, an efficient 2B-parameter hybrid architecture (18 Mamba SSM linear-attention layers + 6 full-attention layers) engineered for high-throughput code retrieval and multimodal embedding.

Compatible with `llama.cpp`, `Ollama`, `LM Studio`, `Unsloth`, `vLLM`, and `sentence-transformers`.

Model Details

Model Description

  • Developed by: Tencent (Base Model); Quantized by ewinregirgojr
  • Model Type: Hybrid Linear-Attention Mamba SSM + Softmax Attention Embedding Model
  • Language(s) (NLP): English, Chinese, Multilingual, and Programming Languages (Python, TypeScript, JavaScript, C++, Rust, Go, SQL, Shell)
  • License: Apache-2.0
  • Base Model: `tencent/WeMM-Embedding-2B`
  • Embedding Dimension: 2,048 dimensions (L2-normalized dense float vector)

Model Sources

  • Base Model Repository: tencent/WeMM-Embedding-2B
  • Quantized Repository: ewinregirgojr/WeMM-Embedding-2B-GGUF
  • Inference Framework: [llama.cpp](https://github…

Source: https://huggingface.co/ewin-reg/WeMM-Embedding-2B-GGUF

Card id model:hf:ewin-reg/WeMM-Embedding-2B-GGUF · collected 2026-10-02 20:59 UTC · JSON