LLMoxie: Exploring Agentic AI for Scientific Software Development
- source_url: https://arxiv.org/abs/2607.02703
- source_type: paper
- platform: arxiv
- author: Landung Setiawan, Anant Mittal, Cordero Core, Anshul Tambay, Carlos Garcia Jurado Suarez, David A. C. Beck, Andrew J. Connolly, Vani Mandava
- original_date: 2026-07-02
- added_date: 2026-07-22
- category: agents
- tags: scientific-software, rse, governance, multi-agent, llmoxie, arxiv
- quality_score: 4
- arxiv_id: 2607.02703
- arxiv_categories: cs.SE, cs.AI, cs.DC, cs.MA
- pdf_url: https://arxiv.org/pdf/2607.02703v1
摘要(中文)
大学 RSE 中心 20 个月实践:三层平台(多云/本地推理、LiteLLM/MLflow 控制面做认证/预算/PII masking/可观测性、编码 agent 应用增强层)+ 开源 RSE-Plugins(Plugin-Agent-Skill,覆盖科学 Python 规范、领域知识、六阶段研究实现工作流与项目生命周期)。核心判断:科研软件看重可引用、可审计、可复现、可扩展,商业 benchmark 优化的 coding agent 不会自动满足。覆盖天文、地球气候、农业、健康等项目。SIGKDD 2026 workshop。
Summary (English)
Describes LLMoxie, an institutional AI platform with three tiers: multi-cloud/on-prem inference; a LiteLLM/MLflow control plane for auth, budgeting, PII masking, and observability; and an application layer for coding agents. An open-source RSE-Plugins ecosystem encodes RSE knowledge as a Plugin-Agent-Skill hierarchy spanning scientific Python practice, domain knowledge, a six-phase research-and-implement workflow, and project lifecycle management. Based on twenty months at a university RSE center across astronomy, earth/climate, agriculture, and health. Argues scientific software is judged by citability, auditability, reproducibility, and extensibility—off-the-shelf coding agents optimized for commercial benchmarks are poorly calibrated. Accepted to ACM SIGKDD 2026 SciSoc Agents workshop.
One-liner
科研软件 Agent 要治理与可审计,不是只会写代码的通用 coding agent。
Source body / metadata
arXiv abstract grounded intake for 2607.02703. PDF: https://arxiv.org/pdf/2607.02703v1
Describes LLMoxie, an institutional AI platform with three tiers: multi-cloud/on-prem inference; a LiteLLM/MLflow control plane for auth, budgeting, PII masking, and observability; and an application layer for coding agents. An open-source RSE-Plugins ecosystem encodes RSE knowledge as a Plugin-Agent-Skill hierarchy spanning scientific Python practice, domain knowledge, a six-phase research-and-implement workflow, and project lifecycle management. Based on twenty months at a university RSE center across astronomy, earth/climate, agriculture, and health. Argues scientific software is judged by citability, auditability, reproducibility, and extensibility—off-the-shelf coding agents optimized for commercial benchmarks are poorly calibrated. Accepted to ACM SIGKDD 2026 SciSoc Agents workshop.
Obsidian evidence
- local_note: OpenClaw定时任务/ClawFeed24小时高价值一览/2026-07-22-ClawFeed24小时高价值一览.md
- intake_run: daily-intake-evening 2026-07-22