LogiShell store Open app

Store › model › Embeddings

WeMM-Embedding-2B-GGUF

by DreamBlooms · source Hugging Face · updated 2026-08-26

other1.5 GB~3 GB RAMsource aliveunlabeled

WeMM-Embedding-2B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 2,048-dimensional L2-…

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:DreamBlooms/WeMM-Embedding-2B-GGUF and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 20:59 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
WeMM-Embedding-2B-BF16.ggufBF164.5 GB
WeMM-Embedding-2B-Q4_K_M.ggufQ4_K_M1.5 GB
WeMM-Embedding-2B-Q8_0.ggufQ8_02.4 GB
mmproj-WeMM-Embedding-2B-BF16.ggufBF16640 MB

From the source README

WeMM-Embedding-2B

WeMM-Embedding-2B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 2,048-dimensional L2-normalized embedding. Audio input is not supported.

Derivation

本仓库为 tencent/WeMM-Embedding-2B 的 GGUF 格式量化版本。

  • 两个量化文件的 GGUF 元数据已注入 `qwen35.pooling_type=3`(last-token pooling),Ollama 可直接识别为 embedding 模型使用。
  • `mmproj-WeMM-Embedding-2B-bf16.gguf` 为独立导出的视觉塔(projector),用于多模态加载。

Installation

pip install torch transformers==5.2.0 "qwen-vl-utils[decord]==0.0.14" \
  "sentence-transformers>=5.7.0" "accelerate>=1.1.0"

Transformers

import torch
from qwen_vl_utils import process_vision_info
from transformers import AutoModel, AutoProcessor

model_id = "tencent/WeMM-Embedding-2B"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModel.from_pretrained(
model_id, trust_remote_code=True, dtype=torch.bfloat16
).cuda().eval()

messages = [{"role": "user", "content": [
{"type": "image", "image": "/path/to/image.jpg"},
{"type": "video", "video": "/path/to/video.mp4"},
{"type": "text", "text": "This can be any text input."},
]}]
text = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=False
)
images, videos, video_kwargs = process_vision_info(
messages,
image_patch_size=16,
return_video_kwargs=True,
return_video_metadata=True,
)
if videos is not None:
videos, video_metadata = zip(*videos)
videos, video_metadata = list(videos), list(video_metadata)
else:
video_metadata = None
inputs = processor(
text=text,
images=images,
videos=videos,
video_met…

Source: https://huggingface.co/DreamBlooms/WeMM-Embedding-2B-GGUF

Card id model:hf:DreamBlooms/WeMM-Embedding-2B-GGUF · collected 2026-10-02 20:59 UTC · JSON