RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust
- ID: a7f15a22
- 原文链接: https://arxiv.org/abs/2609.37916
- 作者: Eugene Hauptmann, Nataliya Kosmyna
- 日期: 2026-09-29
- 更新: N/A
- 分类: infra
- 来源类型: paper
- 标签: tensor-compiler, rust, runtime, multi-backend
- 质量评分: 3/5
- 抓取时间: 2026-10-01T15:57:56+00:00
中文导读
单个 Rust 代码库同时承担编译器与运行时:围绕一套三层 primitive 级 IR 和透明 dispatch contract——每个算子被解析到 native、common-IR 或重写 lowering,无法 legalization 就直接编译报错。同一 IR 目标覆盖 14 种运行时设备(cpu/metal/mlx/ane/cuda/rocm/oneapi/tpu/hexagon/vulkan/webgpu 等)加 Cortex-M INT8 与 FPGA 两条特种代码路径,支持 safetensors/GGUF/ONNX/rten 输入与 F16/BF16/INT4/INT8 量化流,经 TCP/RDMA 做张量/流水线并行扩展。同机同法对比 PyTorch/MLX/IREE/TensorRT/tinygrad 等,all-MiniLM-L6-v2 上 RLX-Metal 每个 batch 最快(batch 32 时 16.6ms vs PyTorch-MPS 26.7ms)。2026 IEEE HPEC 见刊;6 页会议论文,成熟度待验证,但「一份代码管到 Apple ANE 和 Hexagon DSP」的 dispatch 合同设计值得看。
论文信息
- arXiv ID: 2609.37916
- 提交日期: 2026-09-29
- arXiv 分类: cs.DC, cs.AI
- 作者: Eugene Hauptmann, Nataliya Kosmyna
- 链接: https://arxiv.org/abs/2609.37916
Abstract(arXiv 原文)
Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend behavior, deployment guarantees, and performance fallbacks hard to reason about end-to-end. RLX addresses this gap with a single Rust codebase that combines compiler and runtime roles around one primitive-level, three-level intermediate representation (IR), plus a transparent dispatch contract that resolves each operator to native, common-IR, or rewritten lowering and fails compilation when legalization is not possible. The same IR targets fourteen runtime devices (cpu, metal, mlx, ane, cuda, rocm, oneapi, tpu, hexagon, gpu, vulkan, opengl, directx, webgpu) and two specialty codegen paths (Cortex-M INT8 and FPGA), ingests safetensors, GGUF, ONNX, and rten formats, supports F16/BF16/F64/C64 and quantized INT4/INT8 flows with AMP/PTQ/QAT, and scales via tensor-/pipeline-parallel collectives over TCP and RDMA transports. Beyond neural workloads, RLX also extends to scientific/physics-style domains through sparse and dense linear algebra extensions (e.g., CSR LU/CG/matvec and LAPACK- backed factorizations) and 3D Gaussian splatting operators. We evaluate RLX against PyTorch, TensorFlow, JAX, candle, burn, tch, rten, MLX, CoreML, IREE, Glow, TensorRT, and tinygrad under identical input generation and p50 measurement methodology on one host. On all-MiniLM-L6-v2, RLX-Metal is fastest at every batch (e.g., 16.6 ms at batch 32 vs. PyTorch-MPS 26.7 ms). In the MNIST training table, RLX also has the top-throughput entry (graph-fused MLP: 946,487 img/s), above NumPy+BLAS (787,349 img/s), while retaining 100% top-1 parity on reference checks (e.g., Qwen3).
为什么值得关注
「一份代码管到 Apple ANE 和 Hexagon DSP」的 dispatch 合同设计值得看;6 页会议论文,成熟度待验证,当设计参考读。
English Summary
RLX unifies graph compilation and kernel execution in a single Rust codebase built on a primitive-level three-level IR with a transparent dispatch contract that resolves each operator to native, common-IR, or rewritten lowering (failing compilation when legalization is impossible). One IR targets 14 runtime devices (cpu/metal/mlx/ane/cuda/rocm/oneapi/tpu/hexagon/vulkan/webgpu...) plus Cortex-M INT8 and FPGA specialty paths. Against PyTorch/MLX/IREE/TensorRT/tinygrad under identical methodology, RLX-Metal is fastest on all-MiniLM-L6-v2 at every batch (16.6 ms at batch 32 vs 26.7 ms for PyTorch-MPS). IEEE HPEC 2026.