Agent 与自动化 4.0 · 优秀 2026-01-17 · 论文

AEMA: Verifiable Evaluation Framework for Trustworthy and Controlled Agentic LLM Systems

AEMA 针对企业级多 Agent 系统评估:单轮 LLM-as-Judge 或窄 benchmark 难以覆盖多步流程的协调透明和可追溯性论文提出一个过程感知可审计的评估 Agent 框架,负责规划执行并汇总异构工作流评估,同时保留人工监督与 trace record,适合作为 Agent 评测体系的工程参考

打开原文回到归档

AEMA: Verifiable Evaluation Framework for Trustworthy and Controlled Agentic LLM Systems

  • ID: 5a76a55b
  • 原文链接: https://arxiv.org/abs/2601.11903
  • PDF: https://arxiv.org/pdf/2601.11903v1
  • 作者: YenTing Lee, Keerthi Koneru, Zahra Moslemi, Sheethal Kumar, Ramesh Radhakrishnan
  • 日期: 2026-01-17
  • 更新: 2026-01-17
  • 分类: agents
  • 标签: agent-evaluation, multi-agent-systems, llm-as-judge, auditability, trustworthy-ai
  • 质量评分: 4/5
  • 抓取时间: 2026-07-26T12:28:12+08:00

一句话

多 Agent 系统评估需要过程级审计记录,而不是只看单次回答分数

中文导读

AEMA 针对企业级多 Agent 系统评估:单轮 LLM-as-Judge 或窄 benchmark 难以覆盖多步流程的协调透明和可追溯性论文提出一个过程感知可审计的评估 Agent 框架,负责规划执行并汇总异构工作流评估,同时保留人工监督与 trace record,适合作为 Agent 评测体系的工程参考

关键要点

  • 论文备注:Workshop on W51: How Can We Trust and Control Agentic AI? Toward Alignment, Robustness, and Verifiability in Autonomous LLM Agents at AAAI 2026
  • arXiv 分类:cs.AI,适合放入 AAIF 的 agents 频道继续追踪。
  • 内容页由 arXiv 元数据与摘要回填,后续若需要精读可再补充实验细节、威胁模型图和 benchmark 表格。

English Summary

AEMA targets evaluation of LLM-based multi-agent systems, where workflows must show reliable coordination, transparent decision-making, and verifiable performance across changing tasks. The paper criticizes single-response scoring and narrow benchmarks for lacking stability, extensibility, and automation at enterprise multi-agent scale. It proposes Adaptive Evaluation Multi-Agent, a process-aware auditable framework that plans, executes, and aggregates multi-step evaluations across heterogeneous workflows under human oversight, reporting greater stability, human alignment, and traceable records than a single LLM-as-a-Judge.

Abstract

Evaluating large language model (LLM)-based multi-agent systems remains a critical challenge, as these systems must exhibit reliable coordination, transparent decision-making, and verifiable performance across evolving tasks. Existing evaluation approaches often limit themselves to single-response scoring or narrow benchmarks, which lack stability, extensibility, and automation when deployed in enterprise settings at multi-agent scale. We present AEMA (Adaptive Evaluation Multi-Agent), a process-aware and auditable framework that plans, executes, and aggregates multi-step evaluations across heterogeneous agentic workflows under human oversight. Compared to a single LLM-as-a-Judge, AEMA achieves greater stability, human alignment, and traceable records that support accountable automation. Our results on enterprise-style agent workflows simulated using realistic business scenarios demonstrate that AEMA provides a transparent and reproducible pathway toward responsible evaluation of LLM-based multi-agent systems. Keywords Agentic AI, Multi-Agent Systems, Trustworthy AI, Verifiable Evaluation, Human Oversight

Source Metadata

  • arXiv ID: 2601.11903
  • Primary category: cs.AI
  • Categories: cs.AI
  • Comment: Workshop on W51: How Can We Trust and Control Agentic AI? Toward Alignment, Robustness, and Verifiability in Autonomous LLM Agents at AAAI 2026

_Content backfilled by AAIF content-fetcher from OpenCLI arXiv metadata._