The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
- source_url: https://arxiv.org/abs/2607.19292
- source_type: paper
- platform: arxiv
- author: Gjergji Kasneci, Enkelejda Kasneci
- original_date: 2026-07-21
- added_date: 2026-07-23
- category: learning
- tags: ai-safety, socio-technical, governance, instrumentation, memory-poisoning, arxiv
- quality_score: 5
- arxiv_id: 2607.19292
- arxiv_categories: cs.CY, cs.AI, cs.HC
摘要(中文)
观点文:当前 AI 安全话语过度聚焦可见伤害与灾难情景,低估部署中安静可被流程正常化跨组件分散的失败核心不仅是模型是否输出有害内容,而是社会技术系统是否让错误保持可见可争议可遏制可恢复提出五层诊断:认知完整性控制完整性时间完整性组织完整性生态完整性;点名过度依赖检索合法性洗白提示注入奖励黑客记忆投毒评估欺骗虚构人审合成证据污染模型崩塌等,并主张从模型中心评估转向社会技术可靠性arXiv:2607.19292,2026-07-21
Summary (English)
Perspective: AI safety discourse overweights visible harms and catastrophic scenarios while under-instrumenting quieter deployed failures that are plausible, distributed, and normalized by workflows. Central challenge is not only whether a model emits harm, but whether the socio-technical system keeps errors visible, contestable, containable, and recoverable. Five-layer diagnostic: epistemic, control, temporal, organizational, and ecosystem integrity. Highlights under-recognized patternsoverreliance, retrieval legitimacy laundering, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional oversight, synthetic evidence pollution, model collapseand recommends shifting from model-centric eval to socio-technical reliability. arXiv:2607.19292, 2026-07-21.
One-liner
AI 安全不能只盯可见灾难,还要仪表化部署中的静默系统性失败。
Source body / metadata
Fetched via opencli arxiv paper <id> -f json during AAIF content-fetcher backfill. The content note is grounded in the arXiv metadata and abstract.
Abstract
Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that a central safety challenge in modern AI systems is increasingly not only whether a model emits a harmful response, but whether the broader socio-technical system preserves the conditions under which errors remain visible, contestable, containable, and recoverable. We propose a five-layer framework for diagnosing these hidden risks: (1) epistemic integrity, concerning whether evidence and uncertainty are represented honestly enough to support calibrated reliance; (2) control integrity, concerning whether authority, permissions, and action boundaries remain robust under attack and optimization; (3) temporal integrity, concerning whether safety holds across sessions, memory updates, and deployment drift; (4) organizational integrity, concerning whether institutions retain the capacity to audit, assign responsibility, and intervene effectively; and (5) ecosystem integrity, concerning whether AI systems preserve rather than erode the information environment on which future oversight depends. Across these layers, we identify under-recognized risk patterns, including overreliance, uncertainty and legitimacy laundering in retrieval, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollution, and model collapse. We conclude with design and governance recommendations and a research agenda for shifting AI safety from model-centric evaluation toward socio-technical reliability.
Metadata
- arXiv: 2607.19292
- authors: Gjergji Kasneci, Enkelejda Kasneci
- published: 2026-07-21
- categories: cs.CY, cs.AI, cs.HC
- url: https://arxiv.org/abs/2607.19292