Agent 与自动化 4.0 · 优秀 2026-07-09 · 论文

From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

面向企业 LLM agent 产品化,提出 harness 工程模式:把可判定行为下沉到代码/清单/schema/校验,围绕可替换组合边界保证溯源与审计;在韩国五家企业集团公开数据切片上,校验在跨模型替换下仍成立,仅靠 prompt 无法复现合约保证,外部护栏会过拒并掉效用而 harness 保持全效用收录理由:把可审计 agent从口号落到可复用架构与对照实验,对企业落地边界设计很实用

打开原文回到归档

From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

  • Source: https://arxiv.org/abs/2607.08028
  • Platform: arxiv
  • Original Date: 2026-07-09
  • Added: 2026-07-11
  • Category: agents
  • Quality Score: 5
  • Authors: Joongho Ahn , Moonsoo Kim
  • Tags: harness-engineering, enterprise, auditable, contracts, traceability, llm-agent
  • Language: both
  • Fetched: 2026-07-16T20:26:37+08:00

摘要 (Summary)

提出面向企业级 LLM 智能体的 harness engineering 方法,把早期以 prompt 为核心的原型重构为"可追踪可审计"的运行时架构:确定性行为围栏可重建的执行轨迹明确的服务边界与答案契约,支持从原型走向生产环境中的合规部署论文在 4 类企业场景(合规检索源边界实体路由可复现审计)上对所提框架做了系统化评估,展示了相比纯 prompt 工程在可控性与可追溯性上的显著优势

Abstract

Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization adds requirements for source boundaries, entity routing, answer contracts, and reproducible traces. We present a harness-engineering approach that reconstructs this pattern into a traceable, auditable LLM-agent architecture: deterministic behavior moves into code, manifests, schemas, and validation artifacts around a replaceable composition boundary, while source-backed claims remain the authority for runtime answers. We instantiate it on a public-data slice of five Korean corporate groups (25 listed companies) and evaluate three research questions. (1) The harness preserves its source-grounding, entity-routing, trace, output-hygiene, and recommendation-language contracts across the fixed validation scenarios; a fault-injection control confirms the validators flag deliberately broken contracts. (2) The checks the harness enforces held under model substitution: across three hosted models, they passed on all 270 composition-boundary runs; failures were confined to the model-composed side and were caught and recorded. (3) The code-owned guarantees are load-bearing, not reproducible by prompting alone: holding the model fixed and varying only the enforcement layer, prompt instructions alone let recommendation-language and internal-trace-leakage violations reach the reader, which the harness blocks entirely. A bolt-on external guardrail prevents such violations too but over-refuses, dropping utility to 88/120 where the harness preserves full utility (120/120); in this ablation, only code-owned enforcement preserves both safety and utility. The result is a reusable engineering pattern for turning exploratory prototypes into auditable applications with versioned source, control, and validation artifacts.

One-Liner

把企业 LLM 智能体从 prompt 原型重构为可审计 harness:确定性围栏 + 答案契约 + 重建执行轨迹的运行时架构

Full Body

Auto-fetched from the source URL. The body below is the cleaned plain text of the source page (navigation / boilerplate removed).

[2607.08028] From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

Learn more ×

Search

Submit Donate

Log in

Search arXiv

Computer Science > Artificial Intelligence

arXiv:2607.08028 (cs)

[Submitted on 9 Jul 2026]

Title:From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

Authors:Joongho Ahn, Moonsoo Kim View a PDF of the paper titled From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents, by Joongho Ahn and 1 other authors

View PDF HTML (experimental)

Abstract:Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization adds requirements for source boundaries, entity routing, answer contracts, and reproducible traces. We present a harness-engineering approach that reconstructs this pattern into a traceable, auditable LLM-agent architecture: deterministic behavior moves into code, manifests, schemas, and validation artifacts around a replaceable composition boundary, while source-backed claims remain the authority for runtime answers. We instantiate it on a public-data slice of five Korean corporate groups (25 listed companies) and evaluate three research questions. (1) The harness preserves its source-grounding, entity-routing, trace, output-hygiene, and recommendation-language contracts across the fixed validation scenarios; a fault-injection control confirms the validators flag deliberately broken contracts. (2) The checks the harness enforces held under model substitution: across three hosted models, they passed on all 270 composition-boundary runs; failures were confined to the model-composed side and were caught and recorded. (3) The code-owned guarantees are load-bearing, not reproducible by prompting alone: holding the model fixed and varying only the enforcement layer, prompt instructions alone let recommendation-language and internal-trace-leakage violations reach the reader, which the harness blocks entirely. A bolt-on external guardrail prevents such violations too but over-refuses, dropping utility to 88/120 where the harness preserves full utility (120/120); in this ablation, only code-owned enforcement preserves both safety and utility. The result is a reusable engineering pattern for turning exploratory prototypes into auditable applications with versioned source, control, and validation artifacts.

Comments: 32 pages, 6 figures, 16 tables. Reference implementation and evaluation artifacts: this https URL (archived at this https URL)

Subjects:

Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)

ACM classes: I.2.11; D.2.4; D.2.5

Cite as: arXiv:2607.08028 [cs.AI]

(or arXiv:2607.08028v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2607.08028

Focus to learn more

arXiv-issued DOI via DataCite

Submission history From: Joongho Ahn [view email] [v1] Thu, 9 Jul 2026 01:08:33 UTC (1,155 KB)

Full-text links: Access Paper:

View a PDF of the paper titled From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents, by Joongho Ahn and 1 other authors View PDF HTML (experimental) TeX Source

view license

Current browse context:

cs.AI

< prev

| next >

new | recent | 2026-07

Change to browse by:

cs cs.CL cs.SE

References & Citations

NASA ADS Google Scholar

Semantic Scholar

export BibTeX citation

BibTeX formatted citation

×

loading...

Bookmark

Bibliographic Tools