Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
- ID: f2e2d300
- 原文链接: https://arxiv.org/abs/2608.21134
- PDF: https://arxiv.org/pdf/2608.21134v1
- 作者: Luka Ribar, Jeevan Bhoot, Douglas Orr
- 日期: 2026-08-21
- 更新: 2026-08-21
- 分类: models
- 来源类型: paper
- 标签: quantization, vlm, edge-deployment, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-25T15:40:00+00:00
中文导读
把 Llama 3.2 11B Vision 压到 3.7GB 跑在 Arm CPU 上,是这篇最硬的数字。方法组合务实:用模型自己生成训练数据做量化校准(不需要访问原始训练管线),自研 2.7-bit 每参数权重格式专门适配 Arm CPU 推理路径,激活保持 8-bit 避免 VLM 视觉塔精度崩塌。在标准 VQA 任务组上保持强势表现,说明 11B 级 VLM 的端侧部署门槛从「需要旗舰 NPU」降到「CPU 可跑」。端侧多模态选型的必读工程参考。
为什么值得关注
把 Llama 3.2 11B Vision 压到 3.7GB 跑在 Arm CPU 上,是这篇最硬的数字。 实验与数字均来自论文摘要本身。
关键信息
- 论文标题:Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
- 作者:Luka Ribar, Jeevan Bhoot, Douglas Orr
- arXiv:https://arxiv.org/abs/2608.21134
- 发布时间:2026-08-21
- arXiv 分类:cs.CV, cs.LG
English Abstract
Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong performance on a set of standard visual question answering tasks.
English Summary
Quantizes Llama 3.2 11B Vision down to 3.7GB running on Arm CPUs via a self-generated-data calibration pipeline and a custom 2.7-bit per-parameter weight format with 8-bit activations, keeping strong VQA performance - dropping the edge VLM deployment bar from flagship-NPU to CPU-runnable.
Obsidian Notes
- 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成,参考当日论文流水线摘要与评分。
- 中文导读与价值判断均锚定在论文摘要的声明与实验设置上;未补充摘要之外的实验细节。