Introducing Pipette: A benchmarking suite for on-device intelligence
- 原文链接: https://www.liquid.ai/blog/pipette-on-device-ai-benchmarking-by-liquid-ai
- 作者: Liquid AI
- 日期: 2026-08-24
- 分类: infra
- 来源类型: article
- 标签: on-device-llm, benchmark, quantization, mobile, llama-cpp, edge-ai
- 质量评分: 4/5
- 抓取时间: 2026-08-28T23:55:00+08:00
中文导读
Liquid AI 与 Artificial Analysis 联合开源 Pipette 端侧基准平台(Apache 2.0,含 pipette-mgmt/clients/scores 与 iOS/Android 原生客户端):核心论点是「端侧行为是部署系统的属性,不是模型的属性」,评测单元是 model × quantization × runtime × device × context 五维组合,已发布 1,000+ 实验室验证配置、30+ 模型、256-8192 token 上下文,覆盖 M5 Max、iPhone 17 Pro、Galaxy S26 Ultra。公开数据示例:Galaxy S26 Ultra 上 Q4_K_M、256→4096 输入 token,Granite-4.0-H-350M 保持 78.4% decode 吞吐而 Granite-4.0-350M 只剩 33.8%;LFM2.5-8B-A1B decode 比 Qwen3.5-4B 快 2.4 倍但峰值内存仍 5.29 GiB(所有 expert 常驻);iPhone 上 MiniCPM5-1B 快 15.8% 但 LFM 在 MATH-500 高 9 分。Android 路径目前是 CPU-only llama.cpp。端侧选型从此应按目标设备+量化+上下文查表,而非模型卡 server 分数。
为什么值得关注
端云路由与端侧选型的公共基准设施,公开协议+第三方验证,直接改变「看模型卡选端侧模型」的旧做法。
原文(抓取存档·节选)
# Introducing Pipette: A benchmarking suite for on-device intelligence
> 作者: Liquid AI
> 发布时间: 2026-08-24
> 原文链接: https://www.liquid.ai/blog/pipette-on-device-ai-benchmarking-by-liquid-ai
---
Today, in partnership with [Artificial Analysis](http://artificialanalysis.ai/articles/mobile-phone-intelligence-inference), we release **Pipette**, an open-source platform for benchmarking foundation models on edge devices. Pipette is built around a simple premise: on-device behavior is a property of the deployed system, not the model in isolation.
What is being released today:
- **A public dataset of lab-verified results** generated under published reproducibility protocols. It contains five on-device performance metrics for more than 1,000 model × quantization × runtime × device × context configurations. The current data spans 30+ models, multiple quantization formats, llama.cpp builds for macOS, iOS, Windows, and Android, and context lengths from 256 to 8,192 tokens. Initial published results cover MacBook Pro with M5 Max, iPhone 17 Pro, and Galaxy S26 Ultra.
- **Open-source benchmark clients for macOS, Windows, iOS, and Android**, including native iOS and Android apps.
- **An interactive dashboard** connecting quality evaluations with measured throughput, latency, context scaling, and memory use.
- **Apache 2.0-licensed infrastructure**: pipette-mgmt, pipette-clients, and pipette-scores.
## Why On-Device Benchmarks Need Deployment Context
Model cards typically report quality for the original full-precision weights, while on-device deployments often use quantized artifacts. Pipette evaluates quantized artifacts on IFBench, GPQA Diamond, and MATH-500, with FP16 or BF16 results providing a reference where available.
The data shows how strongly these interactions can shape a deployment decision:
**Two 350M models can scale very differently with context.** At Q4_K_M on Galaxy S26 Ultra, Granite-4.0-H-350M retains 78.4% of its decode throughput from 256 to 4,096 input tokens, while Granite-4.0-350M retains only 33.8%.
**Sparse activation can deliver small-model speed without small-model memory use.** At 2,048 input tokens on Galaxy S26 Ultra, LFM2.5-8B-A1B decodes 2.4x faster than Qwen3.5-4B and 2.6x faster than Ministral-3B-Instruct-2512. Despite activating only 1.5B of its 8.5B parameters per token, it still peaks at 5.29 GiB because all expert weights contribute to the model's memory requirements.
**Two similarly sized iPhone models expose a direct speed-quality choice.** At Q4_K_M, MiniCPM5-1B completes the 2,048-in/256-out workload in 3.47 seconds versus 4.12 seconds for LFM2.5-1.2B-Instruct. In quality evaluation of the same Q4_K_M artifacts, LFM scores 9.0 points higher on MATH-500. Neither configuration dominates both axes.
**Nearly identical 8B deployment profiles can hide a task-level reversal.** At Q4_K_M and 2,048 input tokens on M5 Max, Granite-4.1-8B and Ministral-3-8B-Instruct-2512 differ by only 2.4% in decode throughput and 1.2% in peak RAM. In quality evaluations, Granite leads IFBench by 7.3 points while Ministral leads GPQA Diamond by 14.0 points.
Pipette keeps these tradeoffs separate and inspectable rather than collapsing deployment readiness into a single ranking. Performance benchmarks follow published methodology: fixed token shapes, greedy decoding, a discarded warm-up, five measured repetitions, and readiness gating. Evaluations follow a separate published protocol using deterministic, model-blind scoring.