Agent 与自动化 5.0 · 必读 2026-07-21 · 论文

Agents in the Wild: Where Research Meets Deployment

面向把 LLM agentic 系统从研究原型推进到生产的教程,覆盖软件工程科学发现与金融等场景对比学界偏重基准与算法部署侧更关心鲁棒性/安全/可靠性;讨论推理与规划多 Agent 协作与评估,并用制药发现与金融案例提炼成功设计模式与失败缓解(验证流水线回退人在回路)产出侧重设计模式评估清单与跨行业安全可靠部署模板arXiv:2607.19336,2026-07-21

打开原文回到归档

Agents in the Wild: Where Research Meets Deployment

  • source_url: https://arxiv.org/abs/2607.19336
  • source_type: paper
  • platform: arxiv
  • author: Grace Hui Yang, Pranav N. Venkit, Hooman Sedghamiz, Enrico Santus, Victor Dibia, Ioana Baldini
  • original_date: 2026-07-21
  • added_date: 2026-07-23
  • category: agents
  • tags: deployment, tutorial, multi-agent, robustness, evaluation, production, arxiv
  • quality_score: 5
  • arxiv_id: 2607.19336
  • arxiv_categories: cs.AI, cs.CL

摘要(中文)

面向把 LLM agentic 系统从研究原型推进到生产的教程,覆盖软件工程科学发现与金融等场景对比学界偏重基准与算法部署侧更关心鲁棒性/安全/可靠性;讨论推理与规划多 Agent 协作与评估,并用制药发现与金融案例提炼成功设计模式与失败缓解(验证流水线回退人在回路)产出侧重设计模式评估清单与跨行业安全可靠部署模板arXiv:2607.19336,2026-07-21

Summary (English)

Tutorial on agentic LLM systems moving from research prototypes to production in software engineering, scientific discovery, and finance. Contrasts academic benchmark/algorithm focus with deployment needs for robustness, safety, and reliability. Covers reasoning and planning, multi-agent coordination, and evaluation; case studies in pharmaceutical discovery and financial systems surface design patterns and mitigations such as verification pipelines, fallbacks, and human-in-the-loop supervision. Aims to give design patterns, evaluation checklists, and templates for safer industrial deployment. arXiv:2607.19336, 2026-07-21.

One-liner

Agentic LLM 从研究走向生产:关键差距在鲁棒性、安全与可靠部署。

Source body / metadata

Fetched via opencli arxiv paper <id> -f json during AAIF content-fetcher backfill. The content note is grounded in the arXiv metadata and abstract.

Abstract

Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance. While academic work has emphasized benchmarks and algorithmic innovation, deployment raises new challenges around robustness, safety, and reliability. This tutorial brings together researchers and practitioners to explore advances in reasoning and planning, multi agent coordination, and evaluation, highlighting open challenges arising from deployment experience. Through applied case studies in pharmaceutical discovery and financial systems, we analyze common design patterns that make agentic systems successful, and discuss practical mitigation strategies for failure modes, such as verification pipelines, fallback mechanisms, and human in the loop supervision. Attendees will gain a comprehensive view of the field along with concrete design patterns, evaluation checklists, and templates for safe and reliable deployment across industries.

Metadata

  • arXiv: 2607.19336
  • authors: Grace Hui Yang, Pranav N. Venkit, Hooman Sedghamiz, Enrico Santus, Victor Dibia, Ioana Baldini
  • published: 2026-07-21
  • categories: cs.AI, cs.CL
  • url: https://arxiv.org/abs/2607.19336