Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
- ID: f0af643e
- 原文链接: https://arxiv.org/abs/2608.12278
- PDF: https://arxiv.org/pdf/2608.12278v1
- 作者: Avijit Roy, Proma Roy
- 日期: 2026-08-12
- 更新: 2026-08-12
- 分类: learning
- 来源类型: paper
- 标签: underrepresented-languages, tokenization, bengali, fairness, ai-infrastructure, cs.cl, cs.cl-cs.ai-cs.cy
- 质量评分: 4/5
- 抓取时间: 2026-08-14T12:20:00Z
中文导读
论文以孟加拉语为例,指出 AI 教育与语言支持工具虽然常被宣传为"可扩展"的普惠方案,但其底座——训练语料、分词方案、评测基准、部署架构——在模型训练之前就已经系统性地不利于低资源语言使用者。论文识别出四个相互锁定的失败环节:网络存在感缺口(孟加拉语占全球网络内容不足 0.5%,而其人口约占全球 4%)、主要多语言语料中英语与孟加拉语 67:1 的训练 token 赤字、字母音节文字(alphasyllabary)带来的分词惩罚(更高的 token fertility 进一步放大数据赤字)、以及连通性排斥(农村个人互联网渗透率 36.5% 对比城市 71.4%)。作者主张:数据集稀缺应被理解为结构性壁垒而非孤立的技术限制,offline-first 设计应被视为一种公平导向的基础设施策略。
为什么值得关注
把"低资源语言表现差"从模型能力问题重新框定为基础设施与资源分配问题:语料、分词、评测、部署四个环节在训练前就造成结构性不公。对做多语言 AI、评测公平性与边缘市场产品的人来说,这是一份可用于论证资源投入方向的分析框架。
关键信息
- 论文标题:Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
- 作者:Avijit Roy, Proma Roy
- arXiv:https://arxiv.org/abs/2608.12278
- 发布时间:2026-08-12
- arXiv 分类:cs.CL, cs.AI, cs.CY
- 关联标签:underrepresented-languages, tokenization, bengali, fairness, ai-infrastructure
- 论文备注:An associated poster version of this work was presented at the 69th Annual Conference of the International Linguistic Association (ILA 2026), New York, NY, April 30-May 2, 2026
四个结构性失败(论文给出的量化锚点)
1. 网络存在感缺口:孟加拉语占全球网络内容 < 0.5%,人口约占全球 4%。 2. 训练 token 赤字:主要多语言语料中英语:孟加拉语 = 67:1。 3. 分词惩罚:孟加拉语 alphasyllabary 文字带来更高 token fertility,与数据赤字复合叠加。 4. 连通性排斥:农村个人互联网渗透率 36.5%,城市 71.4%(研究聚焦低连通环境下的 AI 辅助教育)。
English Abstract
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world's most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali's alphasyllabary script that compounds the data deficit through higher token fertility; and connectivity exclusion, with individual internet penetration at 36.5% in rural areas compared with 71.4% in urban areas. These failures reflect longstanding resource-allocation decisions, institutional priorities, and design defaults that did not center underrepresented languages in mainstream AI development. We argue that dataset scarcity should be understood as a structural barrier rather than an isolated technical limitation, and that offline-first design should be treated as an equity-oriented infrastructure strategy. We conclude with directions for linguistics and AI research aimed at reducing these structural inequalities.
English Summary
Using Bengali as the case study, the paper argues that AI infrastructure (training corpora, tokenization, benchmarks, deployment architectures) structurally disadvantages underrepresented-language speakers before training even begins. It quantifies four interlocking failures — a <0.5% share of global web content versus ~4% of world population, a 67:1 English-Bengali token deficit in major multilingual corpora, an alphasyllabary tokenization penalty compounding that deficit, and 36.5% rural vs 71.4% urban internet penetration — and reframes dataset scarcity as a structural barrier rather than a technical limitation, advocating offline-first design as an equity-oriented infrastructure strategy.
Obsidian Notes
- 内容由
opencli arxiv paper 2608.12278拉取 arXiv 元数据与摘要生成,正文量化数字均来自论文摘要。 - 中文导读与价值判断锚定在条目已有摘要与论文摘要、作者、日期、分类信息上;未补充论文摘要之外的实验细节。