基础设施 4.0 · 优秀 2026-09-28 · GitHub

Jeff: Jev-compatible 0.8B decision models, trained at home, ~30 ms

Jeff:Qwen3.5 / Gemma 4 微调(0.8B-2B)的零样本分类小模型描述情境并列出选项,单次前向直接返回每个选项的校准概率,不生成文本无需解析,RTX PRO 6000 上约 22msM4 Max (MLX) 约 28ms 一次决策全流程本地硬件训练(单卡 0.8B 约 2 小时),合成数据由开源模型生成;请求格式与 Jev 兼容短微调收益巨大:语音导航任务 held-out 准确率 31.7% -> 95.8%,单 GPU 半小时内小模型做结构化判断大模型做推理的分工再次被验证

打开原文回到归档

Jeff: Jev-compatible 0.8B decision models, trained at home, ~30 ms

原文链接: https://github.com/firelex/jeff
作者: firelex
发布时间: 2026-09-28
源: arXiv外部扫描 (2026-09-29)

摘要

Jeff:Qwen3.5 / Gemma 4 微调(0.8B-2B)的零样本分类小模型描述情境并列出选项,单次前向直接返回每个选项的校准概率,不生成文本无需解析,RTX PRO 6000 上约 22msM4 Max (MLX) 约 28ms 一次决策全流程本地硬件训练(单卡 0.8B 约 2 小时),合成数据由开源模型生成;请求格式与 Jev 兼容短微调收益巨大:语音导航任务 held-out 准确率 31.7% -> 95.8%,单 GPU 半小时内小模型做结构化判断大模型做推理的分工再次被验证

English Summary

Jeff is a set of Qwen3.5 and Gemma 4 fine-tunes (0.8B-2B) for zero-shot classification: describe a situation and list options in plain words, and Jeff returns calibrated probabilities per option from a single forward pass - no generated text, no parsing - about 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max (MLX). Zero-shot options can be anything (support queues, intents, moderation labels, game moves); a short fine-tune on your own examples lifts accuracy much further (31.7% to 95.8% held-out on a voice-navigation task in under half an hour on one GPU). Trained entirely on local hardware with synthetic data written by an open model; uses the same request format as Jev.

为什么值得关注

0.8B 小模型单次前向出校准概率:结构化判断与推理分工的又一实证,本地延迟进入几十毫秒区间

信息源