模型与实验室 4.0 · 优秀 2026-08-21 · 论文

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

把 Llama 3.2 11B Vision 压到 3.7GB 跑在 Arm CPU 上,是这篇最硬的数字方法组合务实:用模型自己生成训练数据做量化校准(不需要访问原始训练管线),自研 2.7-bit 每参数权重格式专门适配 Arm CPU 推理路径,激活保持 8-bit 避免 VLM 视觉塔精度崩塌在标准 VQA 任务组上保持强势表现,说明 11B 级 VLM 的端侧部署门槛从需要旗舰 NPU降到CPU 可跑端侧多模态选型的必读工程参考

打开原文回到归档

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

中文导读

把 Llama 3.2 11B Vision 压到 3.7GB 跑在 Arm CPU 上,是这篇最硬的数字。方法组合务实:用模型自己生成训练数据做量化校准(不需要访问原始训练管线),自研 2.7-bit 每参数权重格式专门适配 Arm CPU 推理路径,激活保持 8-bit 避免 VLM 视觉塔精度崩塌。在标准 VQA 任务组上保持强势表现,说明 11B 级 VLM 的端侧部署门槛从「需要旗舰 NPU」降到「CPU 可跑」。端侧多模态选型的必读工程参考。

为什么值得关注

把 Llama 3.2 11B Vision 压到 3.7GB 跑在 Arm CPU 上,是这篇最硬的数字。 实验与数字均来自论文摘要本身。

关键信息

  • 论文标题:Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
  • 作者:Luka Ribar, Jeevan Bhoot, Douglas Orr
  • arXiv:https://arxiv.org/abs/2608.21134
  • 发布时间:2026-08-21
  • arXiv 分类:cs.CV, cs.LG

English Abstract

Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong performance on a set of standard visual question answering tasks.

English Summary

Quantizes Llama 3.2 11B Vision down to 3.7GB running on Arm CPUs via a self-generated-data calibration pipeline and a custom 2.7-bit per-parameter weight format with 8-bit activations, keeping strong VQA performance - dropping the edge VLM deployment bar from flagship-NPU to CPU-runnable.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成,参考当日论文流水线摘要与评分。
  • 中文导读与价值判断均锚定在论文摘要的声明与实验设置上;未补充摘要之外的实验细节。