【#Tech24H】WeChat announced that its Vision Team has open-sourced the universal multimodal embedding model WeMM-Embedding, which is called over a billion times daily on WeChat’s online services. The release includes three versions: 2B, 4B, and 9B, supporting text, images, video, visual documents, and arbitrary interleaved multimodal inputs. These capabilities are already serving multiple recommendation and search scenarios within WeChat. WeChat has made all model weights, inference code, and evaluation tools publicly available, hoping to help more researchers and developers build multimodal retrieval, recommendation, and agent applications. [ By Zhang Liyan | Tang Ruohan ]


