基础设施 4.0 · 优秀 2026-09-29 · GitHub

Magnitude: self-optimizing inference engine for agents

在 HN 上 Launch 的开源推理引擎:在用户设备上编译并调优 kernel,宣称开源模型比 llama.cpp 最多快 2 倍(Metal decode +92%,CUDA decode +19%),支持 Apple SiliconNVIDIAAMD 与纯 CPU桌面应用内置 magnitude CLI,一键连接 PiOpenCodeHermesCodex 等现有 agent;会话间共享前缀缓存,每 agent 内存低 27%,Apache 2.0,数据不出本机

打开原文回到归档

Magnitude: self-optimizing inference engine for agents

  • ID: 20b52029
  • 原文链接: https://github.com/magnitudedev/magnitude
  • 作者: {'platform': 'github', 'author': 'magnitudedev', 'original_date': '2026-09-29'}
  • 日期: N/A
  • 分类: infra
  • 来源类型: github
  • 标签: inference-engine, local-llm, kernel-tuning, llama-cpp, hn
  • 质量评分: 4/5
  • 抓取时间: 2026-10-01T04:22:00+00:00

中文导读

在 HN 上 Launch 的开源推理引擎:在用户设备上编译并调优 kernel,宣称开源模型比 llama.cpp 最多快 2 倍(Metal decode +92%,CUDA decode +19%),支持 Apple SiliconNVIDIAAMD 与纯 CPU桌面应用内置 magnitude CLI,一键连接 PiOpenCodeHermesCodex 等现有 agent;会话间共享前缀缓存,每 agent 内存低 27%,Apache 2.0,数据不出本机

为什么值得关注

本地自调优推理引擎 Magnitude:设备上编译调优 kernel,宣称比 llama.cpp 最多快 2 倍,一键接入主流 coding agent

Grounded in the fetched README: on-device kernel compilation and tuning, claimed up to 2x faster decode than llama.cpp (+92% Metal, +19% CUDA), cross-session prefix caching with ~27% lower per-agent memory, Apache 2.0, and one-click CLI hookup for existing coding agents.

关键信息

原文摘录

magnitudedev/magnitude: Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD, or just a CPU.

原文链接: https://github.com/magnitudedev/magnitude

Magnitude

#magnitude

Run open models as fast as your hardware allows

Download Magnitude for macOS, Windows, or Linux

⭐ Help us reach more developers and grow the Magnitude community. Star this repo!

demo-9-29.mp4

Get started

#get-started

1. Download Magnitude, install it, and open the app. 2. Choose a recommended model in Discover and download it. 3. Connect your agent in Connections and start using it.

The desktop app includes the magnitude CLI. No separate installation is needed.

Why Magnitude?

#why-magnitude

  • Up to 2x faster than llama.cpp: 92% faster decode on Metal, 19% on CUDA
  • Tuned on your device: kernels are tuned on your hardware before a model runs
  • Built for the best models: hand-optimized kernels for popular open-weight families
  • Memory that flexes: 27% less memory per agent, freed when agents stop
  • Fast concurrent sessions: sessions share prefix caches to prevent slowdown
  • Works with your agent: one click to connect Pi, OpenCode, Hermes, Codex, and more
  • Free, private, open source: no token costs, nothing leaves your machine, Apache 2.0

Up to 2x faster than llama.cpp

#up-to-2x-faster-than-llamacpp

FAQ

#faq

What is Magnitude?

#what-is-magnitude

An open source inference engine that optimizes itself for your hardware. It ships as a desktop app that runs open models and connects them to the agent you already use.

How is it faster than llama.cpp, Ollama, or LM Studio?

#how-is-it-faster-than-llamacpp-ollama-or-lm-studio

They ship kernels precompiled for broad classes of hardware. Magnitude compiles and tunes its kernels on your actual device before a model runs, so they fit your exact chip. See the benchmarks against llama.cpp.

What hardware do I need?

#what-hardware-do-i-need

Any Apple Silicon, NVIDIA, or AMD GPU, or nothing but a CPU. There is no fixed minimum. Smaller machines run smaller models, and more memory lets you run larger ones.

What operating systems does it support?

#what-operating-systems-does-it-support

macOS, Linux, and Windows.

Which models does it support?

#which-models-does-it-support

See the full list at magnitude.dev/models. We write optimized kernels for the most popular open-weight families, which is how we beat generalist engines.

Which agents work with it?

#which-agents-work-with-it

One click connects Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. Anything else works through the OpenAI-compatible API.

Is it private?

#is-it-private

Yes. Prompts, files, and models stay on your machine. No internet needed once a model is downloaded.

English Summary

Open-source inference engine launched on HN that compiles and tunes kernels on the user's device, claiming open models run up to 2x faster than llama.cpp (92% faster decode on Metal, 19% on CUDA) across Apple Silicon, NVIDIA, AMD, and plain CPU. Ships a desktop app with the magnitude CLI, one-click connections to agents like Pi, OpenCode, Hermes, and Codex, shared prefix caches across concurrent sessions, 27% less memory per agent, Apache 2.0, nothing leaves the machine.

Obsidian Notes

  • 内容由 opencli web read 拉取页面正文生成,摘录截断自抓取正文。
  • 中文导读与价值判断锚定在条目摘要与抓取正文上。