基础设施 4.0 · 优秀 2026-08-12 · 文章

Apple Silicon and macOS VMs: 11-16x Faster LLM Inference with llama.cpp

Cua 在 Apple 的 Virtualization.framework 之上做了一层进程级 Metal capability shim,让 macOS guest 里的 llama.cpp 报出 Apple family 9 和 64 KB threadgroup memory,从而选到 SIMD-group matrix / bfloat16 路径,吞吐追到裸金属的 99%M1 Ultra 上 TinyLlama 1.1B Q4_K_M prompt 处理 11.08 倍token 生成 16.36 倍;Gemma 4 12B QAT Q4_0 7.20 倍 / 14.54 倍...

打开原文回到归档

Apple Silicon and macOS VMs: 11-16x Faster LLM Inference with llama.cpp

中文摘要

Cua 在 Apple 的 Virtualization.framework 之上做了一层进程级 Metal capability shim,让 macOS guest 里的 llama.cpp 报出 Apple family 9 和 64 KB threadgroup memory,从而选到 SIMD-group matrix / bfloat16 路径,吞吐追到裸金属的 99%M1 Ultra 上 TinyLlama 1.1B Q4_K_M prompt 处理 11.08 倍token 生成 16.36 倍;Gemma 4 12B QAT Q4_0 7.20 倍 / 14.54 倍;Meta Muse Glimmer 30B Q4_K-M 7.55 倍 / 8.87 倍机制是 paravirtualized GPU 在 guest 里报 Apple 5 家族 + 32 KB threadgroup,llama.cpp 因此走老路径;shim 只动 supportsFamily 和 threadgroup 上限两个返回值,工作仍在 Apple 的虚拟图形路径里执行

English Summary

Cua layers a process-level Metal capability shim atop Apple's Virtualization.framework so llama.cpp inside a macOS guest reports Apple family 9 and 64 KB threadgroup memory, unlocking the SIMD-group matrix / bfloat16 path and recovering ~99% of bare-metal throughput. On M1 Ultra: TinyLlama 1.1B Q4_K_M sees 11.08x prompt / 16.36x token; Gemma 4 12B QAT Q4_0 hits 7.20x / 14.54x; Meta Muse Glimmer 30B Q4_K-M 7.55x / 8.87x. The mechanism is paravirtualized GPU reporting Apple family 5 + 32 KB threadgroup in the guest, forcing llama.cpp onto an older path; only supportsFamily and the threadgroup cap return values are overridden, and execution still rides Apple's virtualized graphics stack.

来源 / Obsidian 引用

  • 本机 Obsidian 同日 digest 落地(ClawFeed 24h / AK-RSS-Digest 89 源 / X-Hot-Brief / 内容选题编排)
  • 评分依据:原文正文(不靠标题/摘要/源声誉)