io_uring without readahead
- ID: 0e7fe806
- 原文链接: https://frn.sh/io-uring
- 作者: Fernando Simes
- 日期: 2026-08-31
- 抓取时间: 2026-09-02T15:31:00Z
- 分类: infra
- 来源类型: article
- 语言: zh
- 标签: io-uring, o-direct, readahead, storage-engine, perf
- 质量评分: 4/5
中文导读
Fernando Simes 8 月 31 日发的测量文,把 Turso 数据库 io_uring + O_DIRECT 后端打开,看为什么 Turso 那条 PR 用应用层 readahead 能跑出明显加速原因拆得很细:不开 readahead 时一次只提交一个 SQE,内核队列里始终只有一个 bio 可合并,所以 SQEs 195,207 个设备侧也接到约 196,000 个合并率 %rrqm 几乎为零;开了 32 页窗口后 SQEs 涨到 218,212,但因为多个 SQE 同时在飞,内核能把相邻扇区拼起来,设备实际只收到约 16,300 个请求%rrqm 91-93%sqqoll 那个 polling 线程在 4 vCPU 机器上吃掉了 65% 的 cycles,关掉 sqqoll 之后 wall time 只差一点点,system time 从 8.46s 降到 1.27s作者最终假设:O_DIRECT 绕开 page cache 时少了一次 CPU copy,顺手把数据带进 L1/L2/L3 的副作用也没了,所以 io_uring 比 syscall 路径多 4.8M cache misses所有做存储引擎或数据库内核的人都该把这篇当 io_uring 调优参考样本
一句话点评
Fernando Simes 8 月 31 日发的测量文,把 Turso 数据库 io_uring + O_DIRECT 后端打开,看为什么 Turso 那条 PR 用应用层 readahead 能跑出明显加速原因拆得很细:不开 readahead 时一次只提交一个 SQE...
English Abstract / Summary
Fernando Simes instruments a Turso database backend running io_uring + O_DIRECT to explain why an application-level readahead patch accelerated it. Without readahead, only one SQE is in flight at a time, the kernel cannot merge bios, and the device sees ~196,000 requests at near-zero %rrqm. With a 32-page readahead window, SQE count rises modestly to 218,212 but the device receives only ~16,300 requests at %rrqm 91-93% because adjacent sectors can now be coalesced. The sqqoll polling thread consumes 65% of cycles on a 4-vCPU box and disabling it cuts system time from 8.46s to 1.27s. Final hypothesis: O_DIRECT eliminates one CPU copy and the cache-warming side effect, costing ~4.8M cache misses versus the syscall path.
Obsidian Notes
- 由
daily-intake-evening2026-09-02 cron 从当日 Obsidian 摘要(论文流水线 / AK-RSS / ClawFeed / X 书签消化)发现并入库存量阶段。 - 中文导读与判断均锚定在条目已有摘要、源页面正文、作者、日期与分类信息;未补充源页面之外的实验细节。