Agent 与自动化 4.0 · 优秀 2026-08-23 · 文章

I built a low-latency AI companion that plays Skyrim with me

作者动手做一个真正陪你玩的 Skyrim AI 伴侣 Varkos,三条目标同时成立:即时且有用(战斗拾取递物跟随多步指令,VR 里尤其不能容忍卡顿)有生命感(个性 + 记住共同经历 + 常开麦克风直接对话,而非从菜单召唤)本地与隐私优先(云 LLM 又贵又慢,单人游戏没必要变成被计费被监视的体验)技术上的亮点是复杂指令处理:条件式指令会注册未来触发器而非立即执行等关联的箭矢命中后再继续计划;计划可以跨步骤保留目标监听进度在世界状态变化时修复或中止,全部无预写脚本工程重心放在延迟与本地推理上,让沉浸感成立时你不是一个人在玩

打开原文回到归档

I built a low-latency AI companion that plays Skyrim with me

Source: https://pantel.is/projects/ai-gaming-companion
Author: Pantelis Kalogiros (@pkalogiros) · Published: 2026-08-23
Type: applied build log (article with video demos)

English Summary

Pantelis Kalogiros built "Varkos", an AI companion for Skyrim that is simultaneously useful, instant, and alive. Three design goals: (1) useful and instant — fight, fetch, loot, carry, follow complex multi-step instructions without feeling buggy, especially in VR where menus break immersion ("it needs to be FAST fast, not just fast"); (2) alive and present — a personality that remembers shared experiences and evolves, with an always-on microphone you talk to rather than a dialogue menu; (3) local and private wherever practical — avoiding metered cloud LLM latency and surveillance in a single-player game.

Notable capabilities demonstrated on video: deferred conditional commands (wait for an arrow fired into the sky, then fetch a potion and bring it), grounded item search over real game state (finds the actual ceremonial sword, or admits it's not there), hide-and-seek as a persistent plan with monitoring and completion conditions, bounded collection plans ("pick up all the items"), and combat with a fast reflex path plus native body control.

The personality layer is the interesting slow path: only starting traits are authored, and the system gradually rewrites both explicit traits and emotional homeostasis from accumulated shared experiences — versioned and reversible. Varkos starts as a demon reincarnated as a dog (mistrustful, proud, sarcastic) and can become domesticated over time. When the game closes he enters "void mode" — blind speech-only contact whose tone depends on how the player has treated him.

The tech stack is the most transferable part:

  • Audio in: always-on mic; Qwen3-ASR 1.7B with custom kernels plus a rolling-partials harness (~40–80ms), Silero/turnpipe VAD, plus lexical turn-completion analysis.
  • Audio out: PocketTTS-Raven (open-sourced) at 20–30ms, with qwen-3-tts for emotionally complex lines; multiple emotion-specific voices pre-loaded in memory.
  • "Thinking": ALE (Action Latent Encoder) — a hybrid of embeddings, small classifiers, explicit rules and traditional ML that detects structure, negation, commands, pronouns, and sequences, matches them against action prototypes including world state, and decomposes plans into linked action slots. Runs in 2–20ms on an M4 MacBook. A local fine-tuned LLM then fuses persona, emotions, history, and the chosen action into the final grounded response; with early prefill the dog can start answering in under 500ms end-to-end.
  • Budget: ~40–80ms ASR + ~20–60ms TTS + ~20ms action analysis + 300–600ms response/grounding.
  • Stance: the author is "bitter-lesson pilled" (big model is better) but argues we've forgotten the lost art of traditional NLP and behavioral graphs; hybrid fast paths win today. Remote providers (e.g. Cerebras-speed gpt-oss-120b / gemma31) don't help much — still weak at holding conversation.

Limitations are honest: local fast models drift over long multi-turn conversations; consumer hardware works today but the design expects 1–2 more years of improvement. The author plans to open-source the ASR harness and eventually the whole system, including a multi-NPC version.

中文摘要(要点)

作者为 Skyrim 造了一只真正陪你玩的 AI 伴侣 Varkos:即时且有用(战斗/拾取/递物/跟随多步指令,VR 里不容忍卡顿),有生命感(个性 + 记住共同经历 + 常开麦克风直接对话),本地优先(不把单机游戏变成按量计费、被监视的体验)。

  • 多步指令不是单次 API 调用:计划可以等待事件、保持目标、监控进度、世界变化时修复或中止;打哪个筒子、找特定剑、捉迷藏,都基于真实游戏状态而非编造。
  • 性格演化走慢路径:只手写初始特征,系统根据共同经历逐步改写显式特征与情绪稳态(可版本化、可回滚);关游戏后进入 void mode,只能靠语音联系,态度取决于你平时怎么对它。
  • 技术栈最值得拆为常开麦克风 + Qwen3-ASR 1.7B 定制内核(40–80ms 转写);PocketTTS-Raven(20–30ms)与 qwen-3-tts 按情绪复杂度分工;核心是 ALE(Action Latent Encoder):嵌入 + 小分类器 + 规则的混合系统,结合世界状态匹配动作原型并分解计划,M4 上 2–20ms 即可返回;再由本地微调 LLM 融合个性与情绪生成回应,端到端可低于 500ms 开口。
  • 作者自述 bitter-lesson pilled(大模型终将更强),但当下靠传统 NLP 与行为图的“失传技话”黑入实时路径;云端推理再快也救不了会话质量。躯诚局限:长对话会漂移;计划开源 ASR harness 与整套系统。

Why It Matters

A rare fully-worked example of local-first sub-second voice agents: instead of asking "can the LLM do it", it asks "what is the cheapest component that can do this slice within the latency budget" — ASR/VAD/TTS/plan decomposition each get a specialized fast path, and the LLM only fuses persona and response. The ALE hybrid (embeddings + tiny classifiers + rules + world-state matching) is a concrete counter-pattern to LLM-planner-everything, and the personality-evolution design (authored seeds + versioned emergent changes) is a reusable pattern for companions beyond games.

Links