资讯资讯

OpenAI 发布 GPT-6 Astra:「欢迎来到 AGI 时代」,API 定价 $10/$50 为上代 2.5 倍OpenAI 发布 GPT-6 Astra:「欢迎来到 AGI 时代」,API 定价 $10/$50 为上代 2.5 倍

📅 2026-09-03 ⏱️ 约 8 分钟阅读⏱️ 8 min read ✍️ AI导航编辑部✍️ AI Nav Editorial 🔗 ai-tokens.cn
AI 导航讯——OpenAI 于当地时间 2026-09-03 发布新一代旗舰模型 GPT-6 Astra(Astra,拉丁语「星辰」),称其为「当今世界智能水平最高、对齐最好的模型」,在计算机操作、软件工程、网络安全、科学研究等领域达到最先进水平;总裁 Brockman 以「欢迎来到 AGI 时代」收尾发布会。API 定价输入 $10、输出 $50 每百万 tokens,为 GPT-5.6 Sol 促销价的 2.5 倍。本文梳理基准成绩、价格拆解、安全争议与开放节奏。OpenAI shipped its new flagship GPT-6 Astra on 2026-09-03 ('astra', Latin for stars), calling it 'the world's most intelligent and aligned model' with state-of-the-art results in computer use, software engineering, cybersecurity and science; President Greg Brockman closed the briefing with 'Welcome to the AGI era'. API pricing is $10/$50 per million tokens — 2.5× GPT-5.6 Sol. Benchmarks, pricing rules, safety caveats and rollout inside.

OpenAI 在当地时间 2026-09-03(北京时间 9 月 4 日凌晨)正式发布新一代旗舰 GPT-6 Astra。官方称其为「当今世界智能水平最高、对齐程度最高的模型」,在计算机操作、网页浏览、软件工程、网络安全、科学研究与专业工作六大领域均达最先进水平。总裁 Greg Brockman 在媒体吹风会上表示:「我认为现在感觉已经进入 AGI 时代并非不合理」,并以「欢迎来到 AGI 时代」收尾。Astra 使用超过 10 万块 GPU 在得州 Stargate 基地完成训练,是 OpenAI 迄今最大规模的训练 run,也是首次由早期模型深度参与监督训练过程的版本。

基准成绩确实是「代际跃迁」量级:ARC-AGI-3 抽象推理从 GPT-5.6 Sol 的 7.8% 跃升至 99.9%(注意为厂商自测口径——ARC Prize 用中立 harness 复测为 62.7%);FrontierMath Tier 4 高阶数学 97.6%;OSWorld 2.0 计算机操作 72.6%(Sol 为 65.7%,平均任务耗时减少约 47%);Terminal-Bench Science 64.6%;SRE-Bench 公开集 88% pass@1;DeepSWE v1.1 74.1%(Sol 70.8%,但 Meta 同周报 Muse Spark 1.3 为 75.4%)。网络安全提升最具冲击力:ExploitBench 漏洞利用 100%(Sol 78.5%),2026 年 6-8 月新漏洞实战成功率 39% vs Sol 的 5.5%——Astra 成为 OpenAI 首个达到内部「Critical」网络安全能力等级的模型,内测中曾发现并利用两个零日漏洞,因此最先进的网络安全能力暂不全面开放。

Agent 与计算机操作是主打场景:可自主填表单、更新 CRM、整理日历、检索资料写摘要、分析科学数据制图、创建网站并跑前端 QA、安装测试软件并按屏幕反馈排障;OpenAI 还展示了将电子原理图转为可制造 PCB、以约人类冠军 4 倍速度完成金融建模世界杯、用 Unity 建游戏场景等案例。Agents' Last Exam 59.3%、ScreenSpot-Pro 92.7%、AutomationBench 41.4%。规格方面:上下文 105 万 tokens(最大输入 922K)、最大输出 12.8 万,支持文本/图像输入,新增异步工具调用与任务执行中调整指令的能力,可直接下达「做一份财务分析并做成演示文稿」这类目标。

价格大幅上涨:API 标准价输入 $10/百万 tokens、输出 $50/百万 tokens,约为 GPT-5.6 Sol 促销价的 2.5 倍,与 Anthropic Fable 5.1 持平;缓存输入 $1/M、缓存写 $12.50/M;Fast 模式速度最高约 2 倍、价格翻倍。⚠️ 长上下文隐藏加价:单请求输入超过 27.2 万 tokens 后,该请求的输入与缓存价格翻倍、输出加价 50%(是整个请求而非超出部分)。这一代没有 Luna/Terra/Sol 系列小模型,只有 Astra 与 Astra Pro 两档。官方称 Astra 在多项评测中 token 用量更省,「重要的是单任务价格」——但第三方普遍认为需实测验证。

安全与争议同样值得注意:对齐表现明显改善——在故意布置的困难任务中越权行为从 Sol 的 48% 降到 0%;但新推理架构把更多计算放进隐藏内部循环,思维链可监控性下降,首席科学家 Pachocki 承认 CoT 监控「脆弱」且趋势「不乐观」。实时监控系统可能减速、暂停甚至终止合法任务(约占 20% 推理算力开销)。批评者指出 OpenAI 未发布自家 AGI 基准 GDPval 的成绩,且 Astra 在部分榜单落后 Claude Fable 5.1。顺带一提:发布当天 ChatGPT/Codex 因北美服务器运营商宕机故障约 2 小时。

开放节奏与建议:即日起先向 Daybreak 安全项目的企业客户开放,未来数日覆盖 ChatGPT Plus/Pro/Business/Enterprise 订阅、官方 API 与 AWS Bedrock;Astra Pro 仅面向 Pro/Business/Enterprise;企业租户默认关闭、需管理员手动开启;免费用户近期无缘;合格 API 客户支持零数据保留(ZDR)。Sam Altman 借势首次明确「一定会做人形机器人」。实操建议:$10/$50 的定价适合高价值、高难度任务(安全审计、科研、复杂 Agent 流程),常规编码与对话继续用 Sol 或更便宜的模型;等第三方复测与价格回落后再决定是否迁移核心工作负载。

OpenAI officially launched GPT-6 Astra on 2026-09-03 (early Sep 4 Beijing time). The company calls it 'the world's most intelligent and aligned model', state-of-the-art across computer use, browsing, software engineering, cybersecurity, science and professional work. President Greg Brockman told the press briefing: 'I think it's not unreasonable to feel that we are now in the AGI era', closing with 'Welcome to the AGI era.' Astra was trained on 100,000+ GPUs at the Stargate site in Texas — OpenAI's largest training run ever, and the first release where earlier models deeply supervised the training process.

The benchmark jumps are generational: ARC-AGI-3 leaps from 7.8% (GPT-5.6 Sol) to 99.9% (vendor-run caveat — ARC Prize's neutral harness scored 62.7%); FrontierMath Tier 4 hits 97.6%; OSWorld 2.0 computer use 72.6% vs Sol's 65.7% with ~47% less average task time; Terminal-Bench Science 64.6%; SRE-Bench 88% pass@1; DeepSWE v1.1 74.1% (Sol 70.8%, though Meta reported 75.4% for Muse Spark 1.3 the same week). Cybersecurity is the sharpest jump: 100% on ExploitBench (Sol 78.5%) and 39% success on June–Aug 2026 fresh vulnerabilities vs 5.5% — making Astra the first OpenAI model rated 'Critical' under its Preparedness framework; it found and exploited two zero-days in internal tests, so the most advanced cyber capabilities are not broadly enabled.

Agentic computer use is the headline scenario: filling web forms, updating CRMs, managing calendars, researching and summarizing, charting scientific data, building sites with frontend QA, installing and debugging software from on-screen feedback; demos include converting schematics to manufacturable PCBs, finishing the Financial Modeling World Cup ~4× faster than the human champion, and Unity scene building. Agents' Last Exam 59.3%, ScreenSpot-Pro 92.7%, AutomationBench 41.4%. Specs: 1.05M-token context (922K max input), 128K max output, text/image inputs, new async tool calls and mid-task instruction changes — you can hand it goals like 'build a financial analysis and turn it into a deck'.

Pricing jumps too: $10 per million input tokens and $50 per million output tokens — 2.5× GPT-5.6 Sol's promotional price, matching Anthropic Fable 5.1; cached input $1/M, cache write $12.50/M; Fast mode up to ~2× speed at 2× price. ⚠️ Long-context surcharge: once a request exceeds 272K input tokens, input and cache prices double and output rises 50% for the entire request. There are no Luna/Terra/Sol siblings this generation — just Astra and Astra Pro. OpenAI says Astra uses fewer tokens on many evaluations and 'price per task is what matters', but third parties say real-world validation is still pending.

Safety and controversy deserve equal attention: alignment improved sharply — goal-unauthorized behavior on deliberately impossible tasks fell from Sol's 48% to 0%; but the new reasoning architecture hides more computation in internal loops, reducing chain-of-thought monitorability, and chief scientist Pachocki admitted CoT monitoring is 'fragile' with an 'unfortunately' negative trend. Real-time monitoring may slow, pause or stop legitimate work (~20% of inference compute). Critics note OpenAI did not publish GDPval — its own AGI yardstick — and Astra trails Claude Fable 5.1 on several indices. Bonus chaos: ChatGPT/Codex went down for ~2 hours on launch day due to a North American server operator outage.

Rollout and advice: available now to enterprises in the Daybreak security program, expanding over the coming days to ChatGPT Plus/Pro/Business/Enterprise, the OpenAI API and AWS Bedrock; Astra Pro is Pro/Business/Enterprise only; enterprise tenants are off by default until an admin enables it; free users get nothing for now; eligible API customers get Zero Data Retention. Sam Altman also confirmed OpenAI 'will definitely' build humanoid robots. Practical take: $10/$50 fits high-value, high-difficulty work (security audits, research, complex agent pipelines); keep routine coding and chat on Sol or cheaper models, and wait for third-party retests and price drops before migrating core workloads.