Introducing Shieldstral.
- ID: 6d7e9807
- 原文链接: https://mistral.ai/news/shieldstral
- 作者 / 日期: Mistral AI | 2026-08-04
- 分类: models
- 来源类型: product
- 标签: safety, guardrails, multimodal, policy-adaptive, open-weights
- 质量评分: 4/5
- 抓取时间: 2026-08-05T15:45:26.663004+00:00
中文导读
Mistral 发布 Shieldstral,一个 3B open-weights multimodal safety classifier,Apache 2.0它把内容审核建模为推理时可输入自然语言 policy 的 yes/no 问答:请求由 InstructQueryDocument 组成,只读取 yes/no logits 并归一化成安全分相比固定 taxonomy guardrail,这种形态更适合不同产品地区和人群切换安全策略;官方称单张 16GB NVIDIA GPU 可运行,并在文本安全拒答检测和多模态审核上达到强表现
为什么值得关注
Shieldstral 的看点是把 guardrail 从固定分类器变成运行时可改 policy 的二元问答模型
English Summary
Shieldstral is a 3B open-weights multimodal safety classifier from Mistral that frames moderation as a policy-adaptive yes/no question at inference time, returning calibrated safety scores without retraining and running on a single 16GB NVIDIA GPU.