Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark
Source: <https://arxiv.org/abs/2609.28449>
Authors: Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, Hadi Hemmati
Published: 2026-09-23
Categories: cs.SE, cs.AI, cs.CL
arXiv: 2609.28449
Abstract
Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mostly limited to snippets or functions. We introduce SWE-Flux, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written manually or judged by LLMs. The benchmark covers singletest and multi-test questions over control flow, loops, program state, dataflow, exceptions, and program invariants. Evaluating five LLMs shows that this task remains challenging. The best model achieves only 37% accuracy. Models perform better on localized behavior such as invariants, intra-procedural control flow, exceptions, and simple loops, but struggle with dataflow, inter-procedural execution, precise state reasoning, and suite-level aggregation. An oracle-harvesting pipeline can generate fresh benchmark variants using input perturbation. It successfully harvests valid variants for almost 90% of the selected instances, and the resulting variants are substantially more challenging for the evaluated models.
Summary
Introduces SWE-Flux, a repository-level benchmark of 480 execution-grounded instances across 12 Python repos with gold answers auto-harvested from instrumented test runs (no LLM judges). Covers control flow, loops, state, dataflow, exceptions, and invariants. Five evaluated LLMs top out at 37% accuracy: they handle invariants, intra-procedural control flow, exceptions, and simple loops but struggle with dataflow, inter-procedural execution, precise state, and suite-level aggregation. An input-perturbation pipeline generates harder variants for ~90% of selected instances.
摘要
推出 SWE-Flux:跨 12 个真实 Python 仓库的 480 个执行驱动实例,金标答案从仪器化测试运行自动采集,不依赖 LLM 评审。覆盖控制流、循环、状态、数据流、异常、不变量。五个评测模型最佳准确率仅 37%:在局部行为上表现较佳,但在数据流、跨过程执行、精准状态与集合级判断上大面紧缺。输入扰动管道可为约 90% 的实例生成更难变体。