基础设施 4.0 · 优秀 2026-07-31 · 文章

Asynchronous I/O in DuckDB: Work, Thread, Work

DuckDB v2.0(2026 秋)将默认对 Parquet/CSV 开异步读:引入 REGULAR + ASYNC 两个线程池,ASYNC 池负责让 fetch 持续在飞worker 拿到 job 即解码,fetch 与 decode 重叠同步读下 row group 的 fetch 未返回时 worker 只能空转;该改造直接面向 EC2/S3 计算存储分离场景(DuckLake + Quack 远程协议),本地 SSD 收益有限对直接查 S3 上 Parquet的湖仓工作流是一次免费吞吐提升

打开原文回到归档

Asynchronous I/O in DuckDB: Work, Thread, Work

  • ID: acc499b6
  • 原文链接: https://duckdb.org/2026/07/31/asynchronous-io
  • 作者: Pedro Holanda
  • 日期: 2026-07-31
  • 分类: infra
  • 来源类型: article
  • 标签: duckdb, async-io, parquet, lakehouse, s3, database
  • 质量评分: 4/5
  • 抓取时间: 2026-08-17T23:51:34+08:00

中文导读

DuckDB v2.0(2026 秋)将默认对 Parquet/CSV 开异步读:引入 REGULAR + ASYNC 两个线程池,ASYNC 池负责让 fetch 持续在飞、worker 拿到 job 即解码,fetch 与 decode 重叠。同步读下 row group 的 fetch 未返回时 worker 只能空转;该改造直接面向 EC2/S3 计算存储分离场景(DuckLake + Quack 远程协议),本地 SSD 收益有限。对“直接查 S3 上 Parquet”的湖仓工作流是一次免费吞吐提升。

为什么值得关注

DuckDB v2.0 异步 I/O 预览:双线程池让 fetch 与 decode 重叠,S3 直查 Parquet 免费提速。

收录理由:数据基础设施的工程化拐点:湖仓直查场景的免费性能提升,附清晰机制图解

关键信息

  • ClawFeed 24小时高价值一览 评分:ClawFeed 评分:8.2/10
  • 来源:ClawFeed 24小时高价值一览(2026-08-17 期)
  • Obsidian 证据:OpenClaw定时任务/ClawFeed24小时高价值一览/2026-08-17-ClawFeed24小时高价值一览.md

原文快照

Asynchronous I/O in DuckDB: Work, Thread, Work

Asynchronous I/O in DuckDB: Work, Thread, Work

Pedro Holanda

2026-07-31 | 21 min

_TL;DR: Starting with v2.0, scheduled for fall 2026, DuckDB will support asynchronous reads of Parquet and CSV files. This can significantly speed up queries when synchronous I/O does not saturate the available bandwidth, as is typical in EC2/S3 compute-storage setups._

It doesn't matter how fast query operators are in a database system if we can't pull in the data quickly. For most of DuckDB's history, however, this problem was largely avoided by pruning data early. By pushing down filters and projections, we could ensure that we only read what we actually needed.

This worked particularly well because DuckDB primarily ran locally, with its main use case being as a quick-draw database engine for querying data directly from your machine's SSD. We could split the data into several partitions, such as row groups for Parquet files or fixed-size buffers for CSV files, and load them with low latency and high bandwidth. As a result, the main bottlenecks were elsewhere: subqueries, joins, aggregations, and so on. The actual data access path received less attention because synchronous access was perfectly suitable for this use case.

As usual, things changed. We realized that DuckDB's architecture was a great fit for querying remotely stored large-scale datasets, such as data lakes (e.g., DuckLake). Since May this year, we can even run DuckDB as a server using the Quack protocol. The original expectation of data files sitting on a local SSD therefore no longer always holds.

The practical implication of these changes is that many current DuckDB setups need to transfer files from remote storage to the machine that will actually process them. For data lakes, for example, a typical setup is to store the data in blob storage, such as S3, and process it on an EC2 machine in the same region. In this setup, latency and bandwidth play a much more significant role. If we cannot issue enough concurrent requests to use the available network bandwidth, performance can suffer drastically, with threads spending a large amount of their time waiting for remote reads instead of processing data.

As an example, let's consider a simple query over a remote Parquet file. For simplicity, let's assume we only have a single thread executing.

FROM read_parquet('s3://bucket/file.parquet');

A Parquet scan is partitioned into row-group-based jobs, with each job containing one or more fetch tasks that issue byte-range requests. With synchronous I/O, the worker thread will be blocked, waiting for the data to arrive at the machine before performing actual work, such as decoding, aggregating, and so on. You can see a visual depiction in the figure

抓取方式:opencli web read(2026-08-17)。完整原文见上方链接。