Google 三周内再推 Gemini 3.8 Flash:介绍价 $0.75/$3.75,Cyber 网络安全变体齐发Google 三周内再推 Gemini 3.8 Flash:介绍价 $0.75/$3.75,Cyber 网络安全变体齐发
Google 在 2026-09-02 一次性发布两款模型:Gemini 3.8 Flash 与 Gemini 3.8 Flash Cyber。这是六周内第三个 Flash 迭代——距 3.7 Flash 仅三周。3.8 Flash 定位「最聪明的干活主力(workhorse)」,API 模型 ID 为 gemini-3.8-flash,已在 Google AI Studio 与 Gemini API 正式可用(GA),可直接上生产。
能力上,3.8 Flash 在长程软件工程基准 DeepSWE v1.1 拿到 73.7%(3.7 Flash 为 65.3%),逼近体量更大的 Claude Opus 5(74.0%);HLE-Verified 多步推理 54.9%;在金融 Agent(Vals Finance Agent V2 61.4%)与 Harvey 法律 Agent 基准上超过 Opus 5、GPT-5.6 等旗舰。关键设计取向是「更勤奋」:面对复杂任务会执行更多推理步骤、迭代调用工具并自查结果。上下文 1M tokens、输出上限 64K,支持文本/图像/视频/音频/PDF 输入。
价格是与 3.7 完全持平的介绍价:输入 $0.75/百万 tokens、输出 $3.75/百万 tokens,有效期至 2026-12-31;2027-01-01 起标准价翻倍至 $1.50/$7.50。缓存输入 $0.075/M(缓存存储 $0.50/M·小时),批处理与 Flex 均 5 折,Priority 档 $1.35/$6.75。注意思考 token 计入输出计费——多家媒体实测其单任务输出 token 约多 30%、实际账单约高 40%。成本敏感场景建议 thinking_level=low,或继续用 3.7 Flash(仍完全支持)。
使用与迁移:思考等级为 low / medium(默认)/ high 三档,旧的数字型 thinking_budget 参数已废弃,须改用 thinking_level 字符串;迁移清单还要求移除 temperature/top_p/top_k、candidate_count 等。渠道覆盖 Google AI Studio、Android Studio、Antigravity(托管 Agent 已默认切换到 3.8 Flash)、Gemini Enterprise;消费者端 Gemini App、搜索 AI Mode、Sheets(AI Pro/Ultra 订阅可用)。
Cyber 变体面向漏洞检测与自动修补,通过新的 Fairwind Program 向政府、关键基础设施与软件维护者等「可信防御者」受限开放。CyberGym 达前沿水平,20 种编程语言漏洞发现成功率超 70%,CWE-Bench 修补 pass@1 为 47.2%(最强商用大模型为 47.8%,属低成本帕累托前沿);Google Chrome 安全团队实测其正确补丁数是最强商用大模型的 2.6 倍,Wiz 内部渗透测试基准召回率提升 7.5–9.7%、成本降低 2.3–5.2 倍。
对开发者和站长的启示:3.8 Flash 是「标价不涨、单任务变贵」的取舍——用更多推理换更高的任务完成率。Artificial Analysis 测算约 $0.58/智能体任务,是同性能档最便宜的模型之一;做 Agent、代码与知识工作的团队值得在 medium 档实测对比 3.7,长线生产项目则要把 2027 年翻倍后的价格一并写进预算。
On 2026-09-02 Google shipped two models at once: Gemini 3.8 Flash and Gemini 3.8 Flash Cyber — the third Flash release in six weeks, only three weeks after 3.7 Flash. 3.8 Flash is billed as the 'most intelligent workhorse', generally available in Google AI Studio and the Gemini API under the ID gemini-3.8-flash, ready for production.
On capability, 3.8 Flash scores 73.7% on the long-horizon software engineering benchmark DeepSWE v1.1 (3.7 Flash: 65.3%), approaching the much larger Claude Opus 5 (74.0%); 54.9% on HLE-Verified; and beats Opus 5 and GPT-5.6 on financial-agent (Vals Finance Agent V2, 61.4%) and Harvey legal-agent benchmarks. The design philosophy is that it 'works harder': extra reasoning steps, iterative tool calls and self-verification on complex tasks. 1M-token context, 64K output, and text/image/video/audio/PDF inputs.
Pricing matches 3.7 Flash's introductory rate: $0.75 per million input tokens and $3.75 per million output tokens through Dec 31, 2026; standard pricing doubles to $1.50/$7.50 from Jan 1, 2027. Cached input is $0.075/M (cache storage $0.50/M·hour); Batch and Flex are half price; Priority is $1.35/$6.75. Thinking tokens bill as output — press tests found ~30% more output tokens per task and roughly 40% higher real-world cost. Use thinking_level=low, or stay on the fully supported 3.7 Flash, when cost matters most.
Usage and migration: three thinking levels (low / medium default / high); the numeric thinking_budget is deprecated in favor of the thinking_level string, and the migration checklist also drops temperature/top_p/top_k and candidate_count. Available via AI Studio, Android Studio, Antigravity (its managed agent now defaults to 3.8 Flash) and Gemini Enterprise; consumers get it in the Gemini app, Search AI Mode and Sheets with AI Pro/Ultra.
The Cyber variant targets vulnerability detection and automated patching, offered to trusted defenders — governments, critical infrastructure, maintainers — through the new Fairwind Program. It reaches frontier level on CyberGym, exceeds 70% success finding vulnerabilities across ~20 languages, and posts 47.2% pass@1 on CWE-Bench patching (best commercial frontier model: 47.8% — a Pareto frontier at far lower cost). Chrome's security team saw 2.6× more correct patches than the best commercial alternative; Wiz reported 7.5–9.7% higher recall on its pentest benchmark at 2.3–5.2× lower cost.
Takeaway for developers and site owners: 3.8 Flash trades a flat sticker price for costlier tasks — more reasoning in exchange for higher task completion. Artificial Analysis pegs it near $0.58 per agentic task, among the cheapest in its performance tier. Teams doing agents, coding and knowledge work should A/B it against 3.7 at medium effort; long-running production projects should budget for the doubled 2027 price.