From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents
- Source: https://arxiv.org/abs/2607.08028
- Platform: arxiv
- Original Date: 2026-07-09
- Added: 2026-07-11
- Category: agents
- Quality Score: 5
- Authors: Joongho Ahn , Moonsoo Kim
- Tags: harness-engineering, enterprise, auditable, contracts, traceability, llm-agent
- Language: both
- Fetched: 2026-07-16T20:26:37+08:00
摘要 (Summary)
提出面向企业级 LLM 智能体的 harness engineering 方法,把早期以 prompt 为核心的原型重构为"可追踪可审计"的运行时架构:确定性行为围栏可重建的执行轨迹明确的服务边界与答案契约,支持从原型走向生产环境中的合规部署论文在 4 类企业场景(合规检索源边界实体路由可复现审计)上对所提框架做了系统化评估,展示了相比纯 prompt 工程在可控性与可追溯性上的显著优势
Abstract
Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization adds requirements for source boundaries, entity routing, answer contracts, and reproducible traces. We present a harness-engineering approach that reconstructs this pattern into a traceable, auditable LLM-agent architecture: deterministic behavior moves into code, manifests, schemas, and validation artifacts around a replaceable composition boundary, while source-backed claims remain the authority for runtime answers. We instantiate it on a public-data slice of five Korean corporate groups (25 listed companies) and evaluate three research questions. (1) The harness preserves its source-grounding, entity-routing, trace, output-hygiene, and recommendation-language contracts across the fixed validation scenarios; a fault-injection control confirms the validators flag deliberately broken contracts. (2) The checks the harness enforces held under model substitution: across three hosted models, they passed on all 270 composition-boundary runs; failures were confined to the model-composed side and were caught and recorded. (3) The code-owned guarantees are load-bearing, not reproducible by prompting alone: holding the model fixed and varying only the enforcement layer, prompt instructions alone let recommendation-language and internal-trace-leakage violations reach the reader, which the harness blocks entirely. A bolt-on external guardrail prevents such violations too but over-refuses, dropping utility to 88/120 where the harness preserves full utility (120/120); in this ablation, only code-owned enforcement preserves both safety and utility. The result is a reusable engineering pattern for turning exploratory prototypes into auditable applications with versioned source, control, and validation artifacts.
One-Liner
把企业 LLM 智能体从 prompt 原型重构为可审计 harness:确定性围栏 + 答案契约 + 重建执行轨迹的运行时架构
Full Body
Auto-fetched from the source URL. The body below is the cleaned plain text of the source page (navigation / boilerplate removed).
[2607.08028] From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents
Learn more ×
Search
Submit Donate
Log in
Search arXiv
Computer Science > Artificial Intelligence
arXiv:2607.08028 (cs)
[Submitted on 9 Jul 2026]
Title:From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents
Authors:Joongho Ahn, Moonsoo Kim View a PDF of the paper titled From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents, by Joongho Ahn and 1 other authors
View PDF HTML (experimental)
Abstract:Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization adds requirements for source boundaries, entity routing, answer contracts, and reproducible traces. We present a harness-engineering approach that reconstructs this pattern into a traceable, auditable LLM-agent architecture: deterministic behavior moves into code, manifests, schemas, and validation artifacts around a replaceable composition boundary, while source-backed claims remain the authority for runtime answers. We instantiate it on a public-data slice of five Korean corporate groups (25 listed companies) and evaluate three research questions. (1) The harness preserves its source-grounding, entity-routing, trace, output-hygiene, and recommendation-language contracts across the fixed validation scenarios; a fault-injection control confirms the validators flag deliberately broken contracts. (2) The checks the harness enforces held under model substitution: across three hosted models, they passed on all 270 composition-boundary runs; failures were confined to the model-composed side and were caught and recorded. (3) The code-owned guarantees are load-bearing, not reproducible by prompting alone: holding the model fixed and varying only the enforcement layer, prompt instructions alone let recommendation-language and internal-trace-leakage violations reach the reader, which the harness blocks entirely. A bolt-on external guardrail prevents such violations too but over-refuses, dropping utility to 88/120 where the harness preserves full utility (120/120); in this ablation, only code-owned enforcement preserves both safety and utility. The result is a reusable engineering pattern for turning exploratory prototypes into auditable applications with versioned source, control, and validation artifacts.
Comments: 32 pages, 6 figures, 16 tables. Reference implementation and evaluation artifacts: this https URL (archived at this https URL)
Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)
ACM classes: I.2.11; D.2.4; D.2.5
Cite as: arXiv:2607.08028 [cs.AI]
(or arXiv:2607.08028v1 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2607.08028
Focus to learn more
arXiv-issued DOI via DataCite
Submission history From: Joongho Ahn [view email] [v1] Thu, 9 Jul 2026 01:08:33 UTC (1,155 KB)
Full-text links: Access Paper:
View a PDF of the paper titled From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents, by Joongho Ahn and 1 other authors View PDF HTML (experimental) TeX Source
view license
Current browse context:
cs.AI
< prev
| next >
new | recent | 2026-07
Change to browse by:
cs cs.CL cs.SE
References & Citations
NASA ADS Google Scholar
Semantic Scholar
export BibTeX citation
BibTeX formatted citation
×
loading...
Bookmark
Bibliographic Tools