type
Post
status
Published
date
Sep 23, 2026 05:00
slug
ai-daily-en-2026-09-23
summary
OpenAI and Anthropic shipped cheaper frontier models on the same day, and the price war is officially on. GPT-6 Sol/Luna cut API prices roughly in half, while Claude Opus 5.5 dropped 20% with a 60% cache-read discount. Xiaomi open-sourced MiMo-V2.6-Pro, a 1T-parameter model trained for about $3M. Al
tags
AI
Daily
Tech Trends
category
AI Tech Report
icon
📰
password
priority
1
📊 Today's Overview
OpenAI and Anthropic shipped cheaper frontier models on the same day, and the price war is officially on. GPT-6 Sol/Luna cut API prices roughly in half, while Claude Opus 5.5 dropped 20% with a 60% cache-read discount. Xiaomi open-sourced MiMo-V2.6-Pro, a 1T-parameter model trained for about $3M. Alibaba unveiled its Zhenwu V900 chip and a 20GW data center roadmap. Meanwhile, Microsoft Research showed 1,024 self-organizing agents beating a single orchestrator.
🔥 Trend Insights
- Frontier pricing collapse: OpenAI halved GPT-6 Sol/Luna prices and Anthropic cut Opus 5.5 by 20% within an hour — per-task cost, not raw capability, is now the battleground.
- Open weights go mainstream: Xiaomi's MiMo-V2.6-Pro (1T-A42B, ~$3M training) is the first top-tier open model from outside the big six, with the full training recipe released.
- Agent harnesses get serious: Microsoft's Agensh scales to 1,024 self-organized agents, and Google's RRSI regularizes harness self-improvement — the scaffolding around frozen models is the new frontier.
🐦 X/Twitter Highlights
📈 热点与趋势
- OpenAI 发布 GPT-6 Sol / Luna,单价减半 - Sol 从 $4/$20 降到 $2/$10,Luna 从 $0.20/$1.20 降到 $0.10/$0.50(每百万输入/输出 token),当天在 ChatGPT Work 与 Codex 面向 Plus/Pro/Business/Enterprise/Edu 上线。Artificial Analysis 测出每任务成本 Sol $1.06(前代 $1.99)、Luna $0.07(前代 $0.18);编码 Agent 指数 Sol 57 升 2 分、Luna 41 降 2 分;AA-Omniscience 幻觉率 Sol 从 92% 降到 60%,但 GDPval-AA v2.1 上 Sol 掉约 100 Elo,主要掉在交付物呈现质量和 rubric 覆盖 @sama @ChatGPT @OpenAIDevs @ArtificialAnlys
- Altman:按 per-task 计价市场里没有竞品 - 称 GPT-6 Sol/Luna 的 per-task 价格市场上无可竞争对象,并说 OpenAI API 的目标是在每个价位都提供最好模型、覆盖文本、代码、图像、视频全模态 @sama @sama
- OpenAI 发布 AI 安全标准提案 - 主张由美国主导,要让实验室外的人对技术走向有真实发言权和判断安全与否的清晰方式;标准应防止权力集中,办法之一是让新公司和开源模型公司能竞争 @sama
- Richard Socher 出新书并创立 Recursive_SI - 《The Eureka Machine》今日出版,主题是全栈科学超级智能:人类知识活地图、物理现实模型、高保真仿真、自主实验室、AI 科学家 agent 集群。8 人团队从 AI 自身的科学做起,再扩到物理、化学与临床前生物 @RichardSocher(Richard Socher,前 Salesforce 首席科学家 / You.com 创始人)
- RAND 报告给出美国超级智能战略两个目标 - 依次为:确保美国地缘优势、确保人类带能动性地存续,结论是最大化可选项 @teortaxesTex(AI 内容博主,转述 RAND 报告)
🔧 工具与产品
- Claude Opus 5.5 发布,运行成本比 Opus 5 低 40% - 属新的 Claude 5.5 家族首款,多数任务达到 Claude Fable 5.1 水平。Perplexity 同日开放给所有 Computer 用户,在其 Wide-And-Deep-Research 评估上优于 Fable 5.1 且成本更低,Opus 5.5 成为 Pro/Max 用户的 "Standard" effort orchestrator @claudeai @AravSrinivas(Aravind Srinivas,Perplexity CEO)
- StepFun 开源 Step Code v0.1.0 - 一个 CLI 编码代理,覆盖读代码、改代码、跑测试到交付全流程;Terminal-Bench 2.1 得 80.9%,自建 Multi-Frame 150 任务长程基准得 73.3%,MIT 许可,另带 StepPage 一键发布静态站 @StepFun_ai(阶跃星辰)
- Kimi 浏览器扩展上线 - 原 Kimi WebBridge,在浏览器侧栏里对话即可让 Kimi 导航网站、填表单;重复任务录一次步骤存为技能,下次自动执行 @Kimi_Moonshot(月之暗面)
- Hamel Husain 与 Shreya Shankar 放出 evals 技能插件 - 让编码 agent 承担大部分评估搭建工作,两人称基于 50 多家 AI 公司经验。文中案例:Ramp 自动收据采集准确率从 35% 升到 83%,Shopify 的 AI 工作流比原前沿模型方案快 2.2 倍、便宜 68%,Cursor 调 Auto Balance 路由后成本降 41% @lennysan(Lenny Rachitsky,Lenny's Newsletter 作者)
⚙️ 技术实践
- vLLM v0.30.0 发布,762 commits / 315 贡献者 - 混合注意力热点路径优化覆盖 Kimi K3 的 KDA、AttnRes、MLA,DeepSeek-V4.1-Flash 的 MXFP8 KV 与异步 Engram,Qwen3.8-Flash-Next 的 QSA/PLE 融合并压低稀疏 GQA 开销;新增 HiSparse 在 sparse-MLA decode 下挂主机层、Model Runner V2 把 EAGLE3 式起草带进流水线并行、双键 Gumbel-max 水印生成与检测 @vllm_project
- vLLM 在 1000+ TPU 上驱动 MiMo-V2.6 310B 全参数 RL - Peano AI 用 JAX 做训练、vLLM 做 rollout,验证阶段与训练器 bitwise 一致;训练器与采样器共用一条 TPU ICI fabric,310B 参数 2 秒内搬完 @vllm_project @peano_ai(Peano AI,TPU 训练平台)
- Perplexity 用提示引导自蒸馏把工具调用失败降 21% - 后训练方法合并拒绝采样微调(RFT)与提示引导自蒸馏:模仿好轨迹,同时显式纠正"整体轨迹成功但工具调用发错"的情况;线上 A/B 测试中,后训练 checkpoint 比早期 checkpoint 工具调用失败率降 21.2% @AravSrinivas @perplexity_ai
- 10 个 Opus 5.5 agent 用 15 小时做出 Lean 验证过的最短路径算法改进 - Vals AI 让十个 agent 设计更快的最短路径算法并在 Lean 中证明,产出 C-HD:对已发表界的经形式化验证的改进 @ValsAI(Vals AI,AI 评测公司)
⭐ Featured Content
GPT-6 Sol/Luna 与 Claude Opus 5.5 同日发布,API 价格战正式开打 | 前沿实验室集体下探定价,与「呼吁放缓」叙事形成反差
OpenAI 发布 GPT-6 Sol 与 Luna,作为 Astra 之后的成本效率梯队,沿用 Astra 训练方法,API 价格较 GPT-5.6 促销价直接砍半(Sol 输入 $4→$2、输出 $20→$10;Luna $0.20→$0.10、$1.20→$0.50)。AutomationBench 显示 Sol 在 xhigh effort 下超过 Claude Opus 5 max effort,单任务成本仅为其 9%;Luna high effort 较前代提升 5.4 个百分点且成本低 58%。约一小时后 Anthropic 发布 Claude Opus 5.5,输入/输出降至 $4/$20(降 20%),缓存读价暴跌 60%——对 90%+ 输入走缓存的长 agentic 对话意义重大,并预告 Sonnet 5.5/Haiku 5.5。Simon Willison 附完整价格对比表并指出 GPT-5.6 十一月还将涨价 25%,实际降幅更大。对做模型选型、成本核算与 agent 工作流的团队,这是本季度最直接的一手定价与能力对比。
GPT-6 prompt caching 升级:cache miss 返回结构化诊断 JSON | agent 长任务成本结构的官方工程手册
OpenAI 为 GPT-6 家族上线改进版 prompt caching:默认命中率更高,30 分钟窗口内复用共享前缀可享最高 90% 输入 token 折扣。新增 Prompt Caching Dashboard 追踪命中率与输入构成,以及 diagnostics 工具——cache miss 会返回结构化 JSON(如 reason=tools_changed、受影响 token 数),直接定位是模型/工具/设置/输入哪一环破坏了复用。三条保缓存实践:用显式 cache breakpoint 选择缓存前缀;GPT-6 上可中途调整 reasoning effort 而不破缓存(追加 configuration_update);工具定义/schema/顺序保持稳定,用 allowed_tools 或 tool_choice=none 替代删除定义,新指令追加到上下文末尾。GitHub Copilot 实测未命中 token 占比降超 50%。
Sources: openai.com
小米 MiMo-V2.6-Pro 开源:1T-A42B、训练成本约 300 万美元 | 非六大厂商首次做出顶级开源权重模型,训练配方全公开
小米发布 MiMo-V2.6-Pro(1T 总参 / 42B 激活)开源权重模型,自称训练成本仅约 300 万美元,同步推出 Flash 与 UltraSpeed(同质量下输出速度最高 20 倍)。技术报告披露 RL 沿三轴扩展:全异步大 batch(每次更新 1,568 样本、最高 1M 上下文、每步 3.5-3.7B tokens)、多任务多 harness 环境混合、以及组内相对比较的 grader compute 以提供更精确的长程奖励。更关键的是小米把环境代码与训练配方全部开源:code、ARVO 漏洞复现、general、webdev、music 五套 recipe 加可组合 mini-harness 配置。对做 agent RL 与开源模型选型的团队,这是一份可直接抄的训练栈。
Sources: latent.space
Hamel evals FAQ:40+ 问答覆盖从最小可行 eval 到 LLM judge 校准 | eval-driven development 的收藏级参考
Hamel Husain 把多年 evals 教学与咨询经验整理成一份 40+ 问答的 FAQ,覆盖从「最小可行 eval 设置」到「LLM judge 与人类不一致怎么办」的全链路。核心观点:坚持二值 pass/fail 而非 1-5 Likert、error analysis 优先于自动化指标、警惕 BERTScore/ROUGE 这类相似度指标、合成数据在长尾场景不可靠、以及「不要为每个失败模式都建自动评估器」。对正在搭 eval 体系的团队,这是一份可直接对照检查的决策清单。
Sources: hamel.dev
让 agent 反复优化 Rust 代码,写出比 state-of-the-art 库更快的实现 | 完整可复用的 agentic 迭代优化工作流
Max Woolf 用 agentic LLM 反复迭代优化 Rust 代码,在禁止 unsafe 的前提下写出比现有 state-of-the-art 库更快的实现,累计加速 2x-20x。方法论:用 criterion 做统计显著性基准、把长 prompt 写成 Markdown 文件(全大写+加粗强调细节)、给 agent 自主迭代权限但设护栏、以 UMAP 从零实现为案例。核心洞察是「优化」并不在主流 agentic 模型的 RLHF 训练目标里,但配合基准与约束后 agent 能持续产出真实加速,且随前沿模型迭代效果递增。
Sources: minimaxir.com
Spotify Home 货架生成系统:LLM 生成「货架假设」+ Semantic ID 落地,serving 时零 LLM 推理 | 生成式推荐在真实产品里的三段式架构
Spotify Research 公开 Home 个性化货架(shelf)生成系统:不再从手工模板库挑选,而是用蒸馏开源 LLM 从用户画像生成自然语言「货架假设」(如 glacial ambient post-rock with orchestral textures),再由生成式检索借 Semantic IDs 把假设落地到真实曲库,最后经 alignment 阶段重选条目并重写标题/副标题保证一致性。全流程离线,serving 时无 LLM 推理。离线评测生成式检索在 Hypothesis-to-Shelf Judge 上 0.71 vs 最强基线 0.56(+27%),alignment 把总分从 0.71 提到 1.27(+78%),标题承诺兑现度 0.66→1.31。线上随机曝光显示专辑场景有竞争力,播客/歌单场景落后。
Sources: daily.dev
阿里发布自研芯片 Zhenwu V900 与 20GW 数据中心路线图 | 中国云厂商全栈 AI 基建加码,Qwen 4 已在训练
阿里 T-Head 发布自研加速器 Zhenwu V900,性能为上一代 M890 的 3 倍,216GB 显存、片间带宽 1.2TB/s,2027 Q1 商用;配套 Panjiu 超节点可扩展至 50 万卡集群。公司目标 2032 年全球数据中心容量超 20GW,支撑其「agentic cloud」战略。硬数据:6 月季度资本开支 677 亿元(同比 +75%),自由现金流 -447 亿元;AI 云与算力收入同比 +44.9% 至 484 亿元,云调整后 EBITA 同比 +133%。Qwen 4 已在训练,Qwen 4.5/5 预计达 5-10 万亿参数。港股当日涨约 3%。
自建 eval 打败公开榜单:35B 模型 95% 成功率反超 120B 的 53% | agent eval harness 的隔离与可复现设计
作者用自建 benchmark 对比 Qwen3.6-35B 与 GPT-OSS-120B 跑自家 coding agent,结果小模型 95% 成功率反超 120B 的 53%,核心论点是「公开榜单和模型规模在自建 eval 面前毫无意义」。文章给出 eval harness 的工程做法:eval 与 agent 完全解耦、通过 CLI 子进程调用、用固定 commit 的 seed repo 保证每次 trial 起点一致、agent 以 git branch 交回结果、用结构化 summary.json 评分而不解析 trace。适合正在给 agent 搭评测层的工程师参考其隔离与可复现设计。
Sources: decodingai.com
🎙️ Podcast Picks
Agent Wars!
📍 Source: AI Daily Brief | ⭐ 3/5 | 🏷️ Agent, Product, Regulation | ⏱️ 00:30:57
The fight over who owns the customer relationship when personal agents make shopping decisions. Meta's Muse passed ChatGPT on the App Store, Amazon blocked its shopping features, and Shopify chose to open up. Also covers Grok 4.7, AI accountability, and cross-lab safety testing.
💡 Why Listen: Good snapshot of the agent commerce land grab. If you care about platform ecosystem battles, this is a quick 30-minute catch-up.
📄 Paper Highlights
Agensh: Scaling Organizational Intelligence to 1,024 Agents
Microsoft Research | 🏷️ Multi-Agent, Agent Framework, Code Agent
Drops the central orchestrator entirely — workers self-assign tasks and merge progress asynchronously. Scaling from 1 to 1,024 agents lifted pandoc test-pass from 33.89% to 55.06%.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Alibaba, Qwen Team | 🏷️ Multimodal, MoE, Agent Framework
A natively omni-modal agentic model with a 1M-token context, plus open-source Qwen-MM-Plugins and Qwen-Live-Harness. Aimed at real production workflows like video editing and long-form translation.
Clarification Is Not Correction: LLMs Fail to Let Go
Meta | 🏷️ Agent Memory, Reasoning, Safety
Names a new failure mode: models commit to one interpretation before ambiguity resolves, then treat later clarification as extra context. Coding tasks are hit hardest — worth a read if you build interactive agents.
🐙 GitHub Trending
Step Code v0.1.0 | Open-source CLI coding agent
StepFun's MIT-licensed CLI agent covers the full loop from reading code to running tests and shipping. Scores 80.9% on Terminal-Bench 2.1 and 73.3% on its own 150-task long-horizon benchmark, with one-click static site publishing via StepPage.
GitHub | ⭐ New | 🗣️ Python | 🏷️ Code Agent, CLI, DevTool
vLLM v0.30.0 | High-throughput LLM inference engine
762 commits from 315 contributors. Adds hybrid-attention hot-path optimizations for Kimi K3, DeepSeek-V4.1-Flash, and Qwen3.8-Flash-Next, plus HiSparse host-layer offload and dual-key Gumbel-max watermarking.
GitHub | ⭐ — | 🗣️ Python | 🏷️ Inference, Serving, Performance