模型与实验室 4.0 · 优秀 2026-08-10 · 文章

Ante: Ghost in your shell

Ante 把 coding agent 做成约 15MB 的 Rust 单文件程序,可接 12 种以上模型服务,也可用内置 llama.cpp 和 GGUF 模型离线运行Terminal-Bench 2.1 完整测试覆盖 89 个任务每项 5 次,DeepSeek V4 Flash 组合取得 368/44582.7%20 个并行 Docker 任务相比 Claude Code 少用 7 倍峰值内存9 倍平均 CPU 和 5 倍磁盘 I/O项目方公布原始 Harbor 记录方便复核

打开原文回到归档

Ante: Ghost in your shell

Source: <https://github.com/AntigmaLabs/ante&gt;
Author: Antigma Labs
Original date: 2026-08-10
Captured: 2026-08-11 (AAIF daily-intake-evening)

中文摘要

Ante 把 coding agent 做成约 15MB 的 Rust 单文件程序,可接 12 种以上模型服务,也可用内置 llama.cpp 和 GGUF 模型离线运行。Terminal-Bench 2.1 完整测试覆盖 89 个任务、每项 5 次,DeepSeek V4 Flash 组合取得 368/445、82.7%。20 个并行 Docker 任务相比 Claude Code 少用 7 倍峰值内存、9 倍平均 CPU 和 5 倍磁盘 I/O。项目方公布原始 Harbor 记录方便复核。

English Summary

Antigma Labs ships Ante as a ~15MB Rust single-binary coding agent: pluggable across 12+ model providers, with a built-in llama.cpp / GGUF path for offline runs. Full Terminal-Bench 2.1 (89 tasks, 5 runs each) puts DeepSeek V4 Flash at 368/445 (82.7%). On 20 parallel Docker tasks, Ante uses ~7x less peak RAM, 9x less average CPU, and 5x less disk I/O than Claude Code. Raw Harbor results are published for replay.

一句话

15MB Rust 单文件 coding agent,资源占用比 Claude Code 低一个数量级

Source Body Excerpt

AntigmaLabs/ante: Ghost in your shell. Ante is a self-contained agent harness with a highly optimized core. It works like Claude Code or Codex, with none of their dependencies or model constraints.

原文链接: https://github.com/AntigmaLabs/ante

https://github.com/AntigmaLabs/ante/releases https://antigma.ai/eval https://docs.antigma.ai/ https://discord.gg/CbAsUR434B https://twitter.com/antigma_labs https://huggingface.co/Antigma https://github.com/AntigmaLabs/ante/blob/main/LICENSE

Ante

#ante

Alpha preview: expect breaking changes and incomplete functionality. macOS and Linux only; on Windows we suggest WSL.

Read this first

#read-this-first

Two things many people ask about:

Where is the source? The core harness currently ships as a prebuilt binary; this repo holds the docs, protocol, SDK, and eval pipeline (details). We are working out a way to ship the source code along with the binary, to address security and privacy concerns first, while taking the time to figure out how open source should work in the agentic era. If you have concerns today, run Ante in a sandbox: it is a single binary with minimal runtime dependencies, built to be easy to deploy in a container or on a remote machine.

Is there telemetry? Yes, and it is opt-out: set ANTE_TELEMETRY=off to disable export entirely. What it sends is anonymous — a random installation label you can delete and re-mint, never your username, hostname, or machine id. The RUST_LOG filter also applies to exported logs, a convenience carried over from the Rust ecosystem. A better UX is in the works. Details →

  • * *

A ghost in your shell. Ante is a self-contained coding agent that lives in your terminal and self-organizes. One ~15MB Rust binary from Antigma Labs, zero runtime dependencies, built to get the most out of any model.

It works like Claude Code or Codex, with none of their dependencies or model constraints. It can also be the optimized core for building your own harness and high-performing assistants.

curl -fsSL https://ante.run/install.sh | bash
ante

Every agent claims to be good. Here are numbers you can check:

🥇 Continuously evaled and evolved, in public

#-continuously-evaled-and-evolved-in-public

Ante runs Terminal-Bench 2.1 continuously under official leaderboard constraints: 89 tasks, 5 trials each. Each result pins the exact build you can download and links the raw Harbor run for independent audit. Latest full run: 82.7% with open-weight DeepSeek V4 Flash 0731 (368/445 trials, Ante 0.preview.71, about $68 of inference). DeepSeek reports the same 82.7 for this model, measured with its unreleased DeepSeek Harness in minimal mode.

Live results → · Methodology →

🪶 A fraction of the footprint

#-a-fraction-of-the-footprint

Ante is hand-written Rust with the heavy parts (Grep, git) embedded in one binary, one process, and local inference handled by a pinned, managed llama.cpp. Across the same 20 parallel tasks in Docker, Ante uses ~7× less peak memory, ~9× less average CPU, and ~5× less disk I/O than Claude Code.

https://github.com/AntigmaLabs/ante/blob/main/docs-site/docs/benchmarks/compare_animated.gif

Raw numbers → · Benchmark details →

🔌 Natively offline

#-natively-offline

Ante's inference engine is a pinned, managed version of llama.cpp. Point it at a GGUF file and the whole loop runs on your machine: no API key, no account, no internet.

ante --offline-model ~/.ante/models/Qwen3.5-9B-Q4_K_M.gguf \
  -p "add error handling to src/main.rs"

Offline mode → · nanochat-rs, a toy engine for study →

  • * *

The three are one design decision. An agent you can verify, afford, and run anywhere is light enough to run by the _thousands_: the substrate for self-organizing intelligence.

See it in action

#see-it-in-action

<table><tbody><tr><td width="50%"><p dir="auto"><strong><a href="https://docs.antigma.ai/usage/models-and-thinking&quot; rel="nofollow">Models, Providers &amp; Thinking</a></strong></p><p dir="auto"><animated-image><a target="_blank" rel="noopener noreferrer" href="https://github.com/AntigmaLabs/ante/blob/main/docs-site/static/assets/cookbook/providers.gif&quot; data-target="animated-image.originalLink"><img src="https://github.com/AntigmaLabs/ante/raw/main/docs-site/static/assets/cookbook/providers.gif&quot; alt="Switching provider, model, and effort with /providers" style="max-width: 100%;" data-target="animated-image.originalImage"></a></animated-image></p></td><td width="50%"><p dir="auto"><strong><a href="https://docs.antigma.ai/usage/providing-context&quot; rel="nofollow">Providing Context: Files &amp; Folders</a></strong></p><p dir="auto"><animated-image><a target="_blank" rel="noopener noreferrer" href="https://github.com/AntigmaLabs/ante/blob/main/docs-site/static/assets/cookbook/files.gif&quot; data-target="animated-image.originalLink"><img src="https://github.com/AntigmaLabs/ante/raw/main/docs-site/static/assets/cookbook/files.gif&quot; alt="Adding file context with @ mentions" style="max-width: 100%;" data-target="animated-image.originalImage"></a></animated-image></p></td></tr><tr><td width="50%"><p dir="auto"><strong><a href="https://docs.antigma.ai/usage/steering&quot; rel="nofollow">Interrupting &amp; Steering</a></strong></p><p dir="auto"><animated-image><a target="_blank" rel="noopener noreferrer" href="https://github.com/AntigmaLabs/ante/blob/main/docs-site/static/assets/cookbook/interrupt.gif&quot; data-target="animated-image.originalLink"><img src="https://github.com/AntigmaLabs/ante/raw/main/docs-site/static/assets/cookbook/interrupt.gif&quot; alt="Interrupting the agent with Escape" style="max-width: 100%;" data-target="animated-image.originalImage"></a></animated-image></p></td><td width="50%"><p dir="auto"><strong><a href="https://docs.antigma.ai/usage/login&quot; rel="nofollow">Subscription Login</a></strong></p><p dir="auto"><animated-image><a target="_blank" rel="noopener noreferrer" href="https://github.com/AntigmaLabs/ante/blob/main/docs-site/static/assets/cookbook/connect.gif&quot; data-target="animated-image.originalLink"><img src="https://github.com/AntigmaLabs/ante/raw/main/docs-site/static/assets/cookbook/connect.gif&quot; alt="Connecting to a provider via /connect" style="max-width: 100%;" data-target="animated-image.originalImage"></a></animated-image></p></td></tr></tbody></table>

See all cookbook guides

Quick Start

#quick-start

Installation

#installation

Ante is a single, self-contained binary with no external dependencies: download and run.

curl -fsSL https://ante.run/install.sh | bash

# Install a specific release channel
curl -fsSL https://ante.run/install.sh | bash -s -- nightly

# Install into a directory already on PATH
curl -fsSL https://ante.run/install.sh | ANTE_INSTALL_DIR=/usr/local/bin bash

Modes

#modes

| Mode | Command | Use it for | | --- | --- | --- | | Interactive TUI | ante | day-to-day work in the terminal | | Headless | ante -p "..." | one-shot tasks, scripts, CI | | Server | ante serve | editor plugins and integrations, over a JSONL protocol | | Gateway | ante gateway | running Ante as a Slack or Discord bot |

Headless examples

#headless-examples

# Fix a bug
ante -p "find and fix the failing test in src/auth"

# Review a diff
git diff | ante -p "review this for security issues"

# Use a different provider
ante --provider openai --model gpt-5.5 -p "refactor the database module"

# Resume a saved session
ante --resume ses_01ARZ3NDEKTSV4RRFFQ69G5FAV -p "now add tests"

# Run fully offline with a local GGUF model
ante --offline-model ~/.ante/models/Qwen3.5-9B-Q4_K_M.gguf \
  -p "add error handling to src/main.rs"

Update Ante

#update-ante

ante update

# One-off update from a different channel
ante update --channel nightly

# Roll back or pin to an exact release
ante update --version v0.preview.71

Beyond the headline numbers

#beyond-the-headline-numbers

  • Zero vendor lock-in: bring your own API key, subscription, or local model. Switch between 12+ providers freely. No account required, not even with us.
  • Multi-agent orchestration: spawn sub-agents and coordinate complex tasks across independent, decentralized, and centralized architectures. See the patterns →
  • Channel integrations: run Ante as a Slack or Discord bot with ante gateway.
  • Extensible: custom skills, sub-agents, MCP, and persistent memory across sessions.

Supported Providers

#supported-providers

Ante works with 12+ providers out of the box:

| Provider | Example Models | | --- | --- | | Anthropic | Claude Sonnet 4.5, Opus 4.6 | | OpenAI | GPT-5 family | | Google Gemini | Gemini 3 family | | Grok (xAI) | Grok 4 | | Open Router | Multiple providers | | Local (GGUF) | Any GGUF model via built-in llama.cpp | | ...and more | Vertex AI, Zai, Antix, OpenAI-compatible |

Configure prov

[...content truncated for storage size]