基础设施 4.0 · 优秀 2026-08-27 · 论文

Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners

在 170 个 Pickle/PyTorch 工件145 个 specimen 家族的受控基准上评估 ModelScanModelAuditFickling 三个 ML 工件安全扫描器,核心区分判断准确率与判断可用性:ModelAudit 对 135 个有标注家族 100% 给出确定结论,Fickling 81.5%,ModelScan 仅 49.6%;但在能给出确定结论的条件下 ModelScan 精准率/召回率/F1 全部 100%ModelScan 分析失败的 48 个恶意家族里,ModelAudit 与 Fickling 都给出符合真值的检测结论:单看 F1 会掩盖覆盖率缺口,工具互补与增量检测覆盖才是部署关键,对做模型供应链扫描的团队是直接的选型证据

打开原文回到归档

Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners

External-scan entry · 20260830 · awesome-ai-field-notes

中文摘要

在 170 个 Pickle/PyTorch 工件、145 个 specimen 家族的受控基准上评估 ModelScan、ModelAudit、Fickling 三个 ML 工件安全扫描器,核心区分判断准确率与判断可用性:ModelAudit 对 135 个有标注家族 100% 给出确定结论,Fickling 81.5%,ModelScan 仅 49.6%;但在能给出确定结论的条件下 ModelScan 精准率/召回率/F1 全部 100%。ModelScan 分析失败的 48 个恶意家族里,ModelAudit 与 Fickling 都给出符合真值的检测。结论:单看 F1 会掩盖覆盖率缺口,工具互补与增量检测覆盖才是部署关键,对做模型供应链扫描的团队是直接的选型证据。

English Abstract

Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases where a scanner yields a usable security judgment. We evaluate ModelScan, ModelAudit, and Fickling using a controlled, artifact-backed benchmark on a synthetic corpus of 170 Pickle and PyTorch focused artifacts across 145 specimen families, 135 of which have binary security ground truth and 10 of which are intentionally malformed without labels. We explicitly distinguish non-N/A coverage, analysis completion, definitive security decisions, non-security findings, and unsupported outcomes. On labeled families, ModelAudit produced definitive security decisions for all 135 families (100%), Fickling for 110 (81.5%), and ModelScan for 67 (49.6%). Conditional on making a definitive judgment, ModelScan achieved 100% precision, recall, and F1. Fickling identified no unique true- positive families beyond those found by the combination of ModelAudit and ModelScan. Furthermore, for the 48 malicious families where ModelScan failed to complete its analysis, both ModelAudit and Fickling generated detections consistent with ground truth. These findings underscore the need to separate judgment accuracy from judgment availability, as well as incremental detection coverage from tool-level redundancy.

注:本文件为 external-scan cron 写入的 source body;如需更深入精读,请由 content-fetcher 任务补充完整正文。