基础设施 4.0 · 优秀 2026-03-19 · 论文

Prompt Control-Flow Integrity: A Priority-Aware Runtime Defense Against Prompt Injection in LLM...

PCFI (Prompt Control-Flow Integrity):把 prompt 视为 system / developer / user / 检索文档的有序组合,在转发到后端 LLM 之前跑三段式中间件(词法启发role-switch 检测层次化策略)实现为 FastAPI gateway,在合成 + 半真实 prompt-injection 评测上:所有带攻击标签的请求都被拦截False Positive Rate 0%中位开销仅 0.04 ms强调"provenance + priority aware"是部署侧实用且轻量的护栏,2026-03-19 提交 cs.CR

打开原文回到归档

Prompt Control-Flow Integrity: A Priority-Aware Runtime Defense Against Prompt Injection in LLM Systems

  • ID: 7d949d83
  • 原文链接: https://arxiv.org/abs/2603.18433
  • PDF: https://arxiv.org/pdf/2603.18433
  • 作者: Md Takrim Ul Alam, Akif Islam, Mohd Ruhul Ameen, Abu Saleh Musa Miah, Jungpil Shin
  • 日期: 2026-03-19
  • 更新: 2026-03-19
  • 分类: infra
  • 来源类型: paper
  • 标签: arxiv, paper, prompt-injection, defense, control-flow
  • 质量评分: 4/5
  • 抓取时间: 2026-09-21T04:30Z

中文导读

  • PCFI 的出发点:现有防御把 prompt 当扁平字符串做临时过滤或静态越狱检测,而 API / RAG 场景里 prompt 本来就是有来源、有优先级的结构——system / developer / user / 检索文档。
  • 实现为 FastAPI gateway,在请求进后端 LLM 之前跑三段中间件:词法启发、role-switch 检测、层次化策略执行。
  • 作者自建的合成 + 半真实注入基准上:所有带攻击标签的请求被拦截、FPR 0%、中位处理开销 0.04 ms。
  • 注意边界:结果是自建 benchmark 而非公开第三方评测,价值在于 provenance/priority-aware 防护的工程形态而非绝对数字。

为什么值得关注

PCFI 把 prompt 视作分层结构,三段式中间件拦截 prompt injection,中位开销 0.04 ms0% FPR

关键信息

  • 论文标题:Prompt Control-Flow Integrity: A Priority-Aware Runtime Defense Against Prompt Injection in LLM Systems
  • 作者:Md Takrim Ul Alam, Akif Islam, Mohd Ruhul Ameen, Abu Saleh Musa Miah, Jungpil Shin
  • arXiv:https://arxiv.org/abs/2603.18433
  • 发布时间:2026-03-19
  • arXiv 分类:cs.CR
  • 关联标签:arxiv, paper, prompt-injection, defense, control-flow

English Abstract

Large language models (LLMs) deployed behind APIs and retrieval-augmented generation (RAG) stacks are vulnerable to prompt injection attacks that may override system policies, subvert intended behavior, and induce unsafe outputs. Existing defenses often treat prompts as flat strings and rely on ad hoc filtering or static jailbreak detection. This paper proposes Prompt Control-Flow Integrity (PCFI), a priority-aware runtime defense that models each request as a structured composition of system, developer, user, and retrieved-document segments. PCFI applies a three-stage middleware pipeline, lexical heuristics, role-switch detection, and hierarchical policy enforcement, before forwarding requests to the backend LLM. We implement PCFI as a FastAPI-based gateway for deployed LLM APIs and evaluate it on a custom benchmark of synthetic and semi-realistic prompt-injection workloads. On the evaluated benchmark suite, PCFI intercepts all attack-labeled requests, maintains a 0% False Positive Rate, and introduces a median processing overhead of only 0.04 ms. These results suggest that provenance- and priority-aware prompt enforcement is a practical and lightweight defense for deployed LLM systems.

English Summary

Large language models (LLMs) deployed behind APIs and retrieval-augmented generation (RAG) stacks are vulnerable to prompt injection attacks that may override system policies, subvert intended behavior, and induce unsafe outputs. Existing defenses often treat prompts as flat strings and rely on ad hoc filtering or static jailbreak detection. This paper proposes Prompt Control-Flow Integrity (PCFI), a priority-aware runtime defense that models each request as a structured composition of system, developer, user, and retrieved-document segments. PCFI applies a three-stage middleware pipeline, lexical heuristics, role-switch detection, and hierarchical policy enforcement, before forwarding requests to the backend LLM. We implement PCFI as a FastAPI-based gateway for deployed LLM APIs and evaluate it on a custom benchmark of synthetic and semi-realistic prompt-injection workloads. On the evaluated benchmark suite, PCFI intercepts all attack-labeled requests, maintains a 0% False Positive Rate, and introduces a median processing overhead of only 0.04 ms. These results suggest that provenance- and priority-aware prompt enforcement is a practical and lightweight defense for deployed LLM systems.

Obsidian Notes

  • 内容由 opencli arxiv paper 2603.18433 -f json 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。