{"v":1,"id":"model:hf:DreamBlooms/WeMM-Embedding-2B-GGUF","slug":"model-dreamblooms-wemm-embedding-2b-gguf","kind":"model","category":"embedding","title":"WeMM-Embedding-2B-GGUF","summary":"WeMM-Embedding-2B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 2,048-dimensional L2-…","source":{"provider":"hf","ref":"DreamBlooms/WeMM-Embedding-2B-GGUF","url":"https://huggingface.co/DreamBlooms/WeMM-Embedding-2B-GGUF","rev":"c896b0570a32d4060bc62faa21b60cea4f49d142","fetchedAt":"2026-10-02T20:59:52.548Z","etag":"W/\"2938-qk0OHb7mdVPAzqsM3QCoUIFX7Os\""},"author":{"name":"DreamBlooms","url":"https://huggingface.co/DreamBlooms"},"license":{"spdx":null,"raw":"other","url":"https://huggingface.co/tencent/WeMM-Embedding-2B/blob/main/LICENSE","open":null,"note":"custom license: read it at the source before installing"},"metrics":{"downloads":2880,"downloadsWeek":327838,"likes":10,"takenAt":"2026-10-02T20:59:52.548Z"},"tags":["transformers","gguf","sentence-transformers","multimodal-embedding","text-embedding","image-embedding","video-embedding","mrl","feature-extraction","zh","en","endpoints_compatible","conversational"],"pipeline":"feature-extraction","links":{"github":"ggml-org/llama.cpp","npm":"node-llama-cpp"},"updatedAt":"2026-08-26T21:26:19.000Z","collectedAt":"2026-10-02T20:59:52.548Z","review":{"numbers":["2,880 downloads on Hugging Face","10 likes","license other","1.5 GB for WeMM-Embedding-2B-Q4_K_M.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:52.548Z","http":200},"description":"# WeMM-Embedding-2B\n\n[](https://huggingface.co/collections/tencent/wemm-embedding)\n[](https://arxiv.org/abs/2608.24053)\n[](https://github.com/Tencent/WeMM-Embedding)\n\nWeMM-Embedding-2B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 2,048-dimensional L2-normalized embedding. Audio input is not supported.\n\n## Derivation\n\n本仓库为 [tencent/WeMM-Embedding-2B](https://huggingface.co/tencent/WeMM-Embedding-2B) 的 GGUF 格式量化版本。\n\n- 两个量化文件的 GGUF 元数据已注入 `qwen35.pooling_type=3`（last-token pooling），Ollama 可直接识别为 embedding 模型使用。\n- `mmproj-WeMM-Embedding-2B-bf16.gguf` 为独立导出的视觉塔（projector），用于多模态加载。\n\n## Installation\n\n```bash\npip install torch transformers==5.2.0 \"qwen-vl-utils[decord]==0.0.14\" \\\n  \"sentence-transformers>=5.7.0\" \"accelerate>=1.1.0\"\n```\n\n## Transformers\n\n```python\nimport torch\nfrom qwen_vl_utils import process_vision_info\nfrom transformers import AutoModel, AutoProcessor\n\nmodel_id = \"tencent/WeMM-Embedding-2B\"\nprocessor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)\nmodel = AutoModel.from_pretrained(\n    model_id, trust_remote_code=True, dtype=torch.bfloat16\n).cuda().eval()\n\nmessages = [{\"role\": \"user\", \"content\": [\n    {\"type\": \"image\", \"image\": \"/path/to/image.jpg\"},\n    {\"type\": \"video\", \"video\": \"/path/to/video.mp4\"},\n    {\"type\": \"text\", \"text\": \"This can be any text input.\"},\n]}]\ntext = processor.apply_chat_template(\n    messages, tokenize=False, add_generation_prompt=False\n)\nimages, videos, video_kwargs = process_vision_info(\n    messages,\n    image_patch_size=16,\n    return_video_kwargs=True,\n    return_video_metadata=True,\n)\nif videos is not None:\n    videos, video_metadata = zip(*videos)\n    videos, video_metadata = list(videos), list(video_metadata)\nelse:\n    video_metadata = None\ninputs = processor(\n    text=text,\n    images=images,\n    videos=videos,\n    video_met…\n\nSource: https://huggingface.co/DreamBlooms/WeMM-Embedding-2B-GGUF","install":{"kind":"model","hfId":"DreamBlooms/WeMM-Embedding-2B-GGUF","gated":false,"format":"gguf","files":[{"name":"WeMM-Embedding-2B-BF16.gguf","size":4790841792,"quant":"BF16","sha256":"7e6ec7e7543de6d3880470996254f170d244d21e433b9cd4e2eaddff1da0a9e5"},{"name":"WeMM-Embedding-2B-Q4_K_M.gguf","size":1559772320,"quant":"Q4_K_M","sha256":"72e6fffa943c5272987e5f3fe1d4c5431a30aaba419228e2669d8e5943fd2d92"},{"name":"WeMM-Embedding-2B-Q8_0.gguf","size":2551300032,"quant":"Q8_0","sha256":"441c18c67d1a7a07fb903a7a075062adb64f38c4b87126e2476a4b4d6a178248"},{"name":"mmproj-WeMM-Embedding-2B-BF16.gguf","size":671373120,"quant":"BF16","sha256":"c7271c975cab83ebd12ce2807e7b77e91134afd9f1658bd4dca986212a60560d"}],"totalBytes":9573287264,"suggestedFile":"WeMM-Embedding-2B-Q4_K_M.gguf","requirements":{"ramGb":3,"diskBytes":1559772320,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:DreamBlooms/WeMM-Embedding-2B-GGUF"}}