来源核验:Liquid AI 官方博客 2026-08-12「LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge」/ Hugging Face 模型卡 LiquidAI/LFM2.5-VL-3B / Hugging Face 博客同日发布 / Liquid AI Docs 视觉能力 / Developers Digest 2026-08-12 深度报道 / WebGPU 浏览器 demo / M5 Max 228 tok/s 实测 / Galaxy S26 Ultra 20 tok/s 实测
3.1B 参数的视觉语言模型,让 8B Gemma 输在屏幕理解
8 月 12 日,Liquid AI 放出了一个开源视觉语言模型:LFM2.5-VL-3B。
它的数字一旦摆出来,会让人先停顿两秒再相信:
- 3.1B 总参数(开源权重)
- ScreenSpot-v2 跨桌面/移动/Web 平均 80.7——比 8B Gemma-4-E4B-it 的 51.2高 29.5 分
- RefCOCO 定位精度 87.9——比上一代 LFM2-VL-3B 的 57.1高 30.8 分
- ToolSandbox 26.4 → 59.5