Can You Check That? The Checkability Boundary for Local LLM Network Automation
原文链接: https://arxiv.org/abs/2609.31540
作者: Maleeha Masood, Momina Nofal
发布时间: 2026-09-25
源: arXiv外部扫描 (2026-09-29)
摘要
arXiv 2609.31540 提出 checkability 判据:网络自动化任务若存在廉价确定性的内在检查(能拒绝违反必要正确性条件的输出),就适合本地小模型推理落地为 Touchstone 流水线7 个 1-8B SLM 生成候选任务内在检查拒绝未决输入升级给 frontier LLM冲突检测和意图翻译分别达到 98.6% / 93.8% 端到端准确率,只需升级 16% / 17% 的输入;没有内在检查的纯知识任务(TeleQnA)则追不上 frontier 基线部署规则:能精确廉价检查的留本地,其余升级
English Summary
Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for direct use. This work introduces checkability as a criterion for determining which tasks are suitable for local inference. A task is checkable when it exposes a cheap, deterministic test - an intrinsic check - that rejects outputs violating a necessary correctness condition. We instantiate this idea in Touchstone, a local-first pipeline that uses seven off-the-shelf SLMs (1-8B parameters) to generate candidates, uses task-specific intrinsic checks to reject responses, and escalates unresolved inputs to a frontier LLM....
为什么值得关注
扩展 AAIF 对应主题线