据腾讯研究院AI速递,谷歌10月6日开源EmbeddingGemma 2。这款740M参数的嵌入模型是其首个原生多模态开源嵌入模型,可将文字、代码、图片、视频、音频映射进同一个768维向量空间,支持100多种语言,上下文长度8K、为上一代的4倍。
基准表现上,其图像检索得57.3分、视频检索50.7分,对比同类开源模型Jina v5 Omni-Nano的31.6分与31.2分优势明显;代码检索MTEB Code从上一代的68.76升至78.68。架构上,文本270M、视觉170M、音频300M三个模块可按需加载;在Pixel 11 Pro上量化后完整模型约占567MB内存,纯文本场景最低约191MB。llama.cpp已提供支持,可与Gemma 4搭配做离线RAG。
【AICOR 点评】嵌入模型是检索增强与企业知识库的地基工程,它的开源与端侧化,意味着中小企业能以极低成本自建覆盖文档、图像、音视频的全格式企业记忆,数据不出内网。多模态统一向量空间的价值在于跨格式关联检索——把图纸、会议录音与工单放进同一个搜索框,这才是知识管理从存档走向调用的分水岭。
Copyright Notice This article is AI-assisted, rewritten from public reports. Copyright of the information belongs to the original authors and media; content is for industry sharing only and does not constitute investment or business advice.
For copyright concerns, contact AICOR (400-601-8080 / WeChat: aicor-ai); we will handle it promptly upon notification.