Apple Silicon and macOS VMs: 11-16x Faster LLM Inference with llama.cpp
- ID: 800ebb0d
- Original: https://github.com/trycua/cua/blob/main/blog/gpu-passthrough-macos-vms.md
- Added: 2026-08-12
- Source: github / trycua
- Original Date: 2026-08-12
- AAIF Category: infra
- Quality Score: 4
- Status: active
- Source Type: article
- Language: en
- Tags: apple-silicon, llama-cpp, inference, macos-vm, metal
中文摘要
Cua 在 Apple 的 Virtualization.framework 之上做了一层进程级 Metal capability shim,让 macOS guest 里的 llama.cpp 报出 Apple family 9 和 64 KB threadgroup memory,从而选到 SIMD-group matrix / bfloat16 路径,吞吐追到裸金属的 99%M1 Ultra 上 TinyLlama 1.1B Q4_K_M prompt 处理 11.08 倍token 生成 16.36 倍;Gemma 4 12B QAT Q4_0 7.20 倍 / 14.54 倍;Meta Muse Glimmer 30B Q4_K-M 7.55 倍 / 8.87 倍机制是 paravirtualized GPU 在 guest 里报 Apple 5 家族 + 32 KB threadgroup,llama.cpp 因此走老路径;shim 只动 supportsFamily 和 threadgroup 上限两个返回值,工作仍在 Apple 的虚拟图形路径里执行
English Summary
Cua layers a process-level Metal capability shim atop Apple's Virtualization.framework so llama.cpp inside a macOS guest reports Apple family 9 and 64 KB threadgroup memory, unlocking the SIMD-group matrix / bfloat16 path and recovering ~99% of bare-metal throughput. On M1 Ultra: TinyLlama 1.1B Q4_K_M sees 11.08x prompt / 16.36x token; Gemma 4 12B QAT Q4_0 hits 7.20x / 14.54x; Meta Muse Glimmer 30B Q4_K-M 7.55x / 8.87x. The mechanism is paravirtualized GPU reporting Apple family 5 + 32 KB threadgroup in the guest, forcing llama.cpp onto an older path; only supportsFamily and the threadgroup cap return values are overridden, and execution still rides Apple's virtualized graphics stack.
来源 / Obsidian 引用
- 本机 Obsidian 同日 digest 落地(ClawFeed 24h / AK-RSS-Digest 89 源 / X-Hot-Brief / 内容选题编排)
- 评分依据:原文正文(不靠标题/摘要/源声誉)