AI Tech Daily - 2026-09-25
2026-9-25
| 2026-9-25
字数 4163阅读时长≈ 11 分钟
type
Post
status
Published
date
Sep 25, 2026 05:00
slug
ai-daily-en-2026-09-25
summary
The AI infrastructure race got a geopolitical twist: the White House is reportedly telling OpenAI and Anthropic to hold new models from UK testers until US review, while Google literally sends TPUs to orbit with Project Suncatcher launching October 1. On the cost front, Vercel's AI Gateway shows Ant
tags
AI
Daily
Tech Trends
category
AI Tech Report
icon
📰
password
priority
1

📊 Today's Overview

The AI infrastructure race got a geopolitical twist: the White House is reportedly telling OpenAI and Anthropic to hold new models from UK testers until US review, while Google literally sends TPUs to orbit with Project Suncatcher launching October 1. On the cost front, Vercel's AI Gateway shows Anthropic's spend share collapsing from 69% to 40% in two months as OpenAI, Kimi K3 and DeepSeek eat the difference. Meanwhile GitHub Security Lab open-sourced an LLM fuzzing agent, and Nubank's simulation-first CX agent workflow lifted tNPS by 36.69 points in live A/B tests.

🔥 Trend Insights

  • Agent cost governance goes mainstream: Vercel's gateway data shows model spend shifting fast, and Accenture's new router paper recovers 14-21% of enterprise model spend — routing is now a first-class infra problem.
  • Sandboxing converges across stacks: Homebrew 7.0.0 and coding agents like Claude Code and Codex now share the same kernel primitives (Landlock, Seatbelt, seccomp), because both run untrusted code.
  • Simulation before deployment: Nubank screened 16,000+ synthetic conversations before shipping, and GitHub's Taskflow does the same for fuzzing — testing agents in silico is becoming standard practice.

🐦 X/Twitter Highlights

📈 热点与趋势

  • Vercel AI Gateway:Anthropic 支出占比两月内从 69% 掉到 40% - OpenAI 同期从 10% 升到 24%,token 数已居第一,图像生成占 62%;Kimi K3 与 DeepSeek 吃掉 Anthropic 流失份额的约一半,Opus 5.5 两天内占到 10% @rauchg(Guillermo Rauch,Vercel CEO)
  • Google 把 TPU 送上天,Project Suncatcher 首星 10 月 1 日发射 - 搭载 SpaceX Transporter-18 从范登堡升空,卫星由 Google 与 Planet 合作建造,目的是验证 TPU 能否在轨存活并运行 @sundarpichai(Sundar Pichai,Google CEO)@cb_doge(科技内容账号)
  • 宇树 19 台全尺寸人形机器人与 120 名舞者同台 - 9 月 22 日上海世界技能大赛开幕式,现场观众超 1 万人,机器人集群为 AI 自主驱动并全球实时直播 @UnitreeRobotics(宇树科技)
  • Simon Willison:coding agent 让软件工程更难,不是更简单 - 称用好它们需要极高纪律与知识储备;Gergely Orosz 跟帖认为非开发者写生产代码短期内不会发生 @simonw(Datasette 作者 / 独立开发者)

🔧 工具与产品

  • Perplexity 发布 Fast Search,Rust 引擎 Photon 做到 p95 230ms - 单次搜索 p50 160ms、p95 230ms,官方称 95% 结果在 230ms 内返回;引擎由小团队加数百个自动研究循环搭出,现已是 Nous Portal 订阅者 Hermes Agent 的默认搜索、全部档位免费 @AravSrinivas(Aravind Srinivas,Perplexity CEO)@NousResearch(Nous Research,开源 AI 研究组织)
  • Scenario 开源 GameDev OS:给编码 agent 9 个专家、64 个技能 - 含 2D/3D 美术、环境、美术总监、音效、视频、营销、技术总监与成本管理员,各自绑定 Seedance、GPT、Meshy、Rodin 等模型,可串联出图→建模→场景→音效→预告片 @emmanuel_2m(Emm,Scenario.com)
  • Odyssey 发布 Agora-2 多智能体世界模型 - 支持最多 20 个人类与 agent 在同一环境实时交互模拟,多人研究预览已开放 @odysseyml(Odyssey,世界模型初创公司)
  • OrcaSAQ-2 27B:体积缩 78.3%,SWE-bench Verified 70.0 - 混合精度 Qwen3.8,55.59GB 压到 12.06GB、3.21 bpw,Top-1 一致率 93.2%,Terminal-Bench 2.1 得 58.4,262K 上下文,面向编码、终端、浏览器与安全 agent @OrcaRouter(OrcaRouter,模型量化路由服务)
  • Muse 给每位用户一台云电脑,用 Secure VM + Sentinel 防提示注入 - runtime cell 独立 root 文件系统(完整 Ubuntu 镜像),敏感操作与凭据都存在 cell 之外由 Sentinel 监管;文件在 Library 标签页可直接浏览,Settings 里可下载 agent 数据 @dps(David Singleton,Meta 工程师)
  • Quail:开源 AI-SQL 引擎,单张 H100 跑到 1B+ 输入 token/分钟 - 与 Modal 合作,把查询规划与 LLM 推理放在一起调度,针对 AI 数据算子的新推理负载 @sh_reya(Shreya Shankar,UC Berkeley 研究者)

⚙️ 技术实践

  • vLLM 与 RLKernel 在 ROCm 上做到训练/rollout logprob 逐位一致 - 8× AMD MI300X 跑 200 步 Qwen3-8B GRPO,Megatron 训练与 vLLM rollout 零 mismatch,严格路径对齐了归约顺序、中间精度、舍入点与数学原语 @vllm_project(vLLM,开源推理引擎)
  • GLM-5.3 在 8× AMD MI355X 上单用户解码 469 tok/s - vLLM 负责 prefill、TileRT 通过 V1 connector 接管延迟敏感的 decode,FP8 下比 GB300 上的 TRTLLM FP4 快 40% 以上 @vllm_project @SemiAnalysis_(SemiAnalysis,半导体分析机构)
  • Meta AI 给 agent 配专职记忆 agent,Claude Sonnet 4.5 基准从 37.6% 升到 45.9% - 主行动 agent 配记忆 agent:结构化记录工具与历史错误,只在关键节点注入短提醒,对抗长上下文里的"忘掉早前错误" @DeepLearningAI(DeepLearning.AI,AI 教育机构)
  • MiniMax-H3 视频模型上 AMD MI355X,5 秒视频 1.3 秒生成 - Nunchux 做的推理优化,8 卡下比 SGLang 快最多 26.7 倍,支持流式生成中改提示词 @MiniMax_AI(MiniMax)
  • oMLX 0.7.0rc1:Qwen3.8-27B 解码 +131%,下轮 prefill 从 1174 token 降到 37 - M5 Max 128GB 上 Qwen3.8-Flash-Next prefill 提升 32%,新增部分块缓存(无需重算未满缓存块的 token)、Ternary Bonsai 2、MiMo V2.6 多模态理解与 MCDMA RDMA 支持 @jundotkim(Jun Kim,oMLX 作者)
  • Claude Code 用 /advisor 把 Fable 挂成旁路顾问 - 主会话由 Opus 5.5 high 跑,三个 medium 子代理分担读码、改测、查文档,Fable 只在定计划前、同一错误复发、宣布完成前介入;附完整子代理与 effortLevel 配置提示词 @Voxyz_ai(Vox,AI 工具博主)

⭐ Featured Content

GitHub 开源 AI 模糊测试 Taskflow:LLM 决策 + MCP 执行的责任分离范式 | 把 C/C++ fuzzing 全流程交给 agent 的完整参考实现
GitHub Security Lab 开源 Fuzzing Taskflow,让 LLM agent 端到端跑完模糊测试:识别入口点、分析构建系统、写 harness、跑 AFL++、读覆盖率、迭代改进、triage 崩溃并生成漏洞报告。架构上刻意做责任分离——agent 只做决策,MCP 工具只暴露 run_afl_for / compile_harness 等原语,状态全部走 SQLite;每个 harness 构建 .afl 与 .cov 两份二进制形成 coverage-feedback 闭环。文中明确警告该流程会在宿主机直接执行 LLM 选定的构建命令,须在一次性环境运行。对做 agentic pipeline 的团队,这是一份可直接 codespace 跑通的三层架构(shell driver + taskflow YAML + MCP 工具)范本。
来源:github.blog
「chat 是错的 UI」:GitHub 提出让 agent 自己造工具,后续交互免费 | canvas 交互范式与 token 经济学
GitHub 官方博客抛出反直觉观点:chat 是 LLM 的默认界面,但大多数任务场景下它是错误的 UI。作者以 Copilot app 的 canvas 为例——运行在 app 内、无浏览器外壳的全栈小应用,可与 agent 双向通信、调第三方 API、本地执行代码。核心洞察是:与其让 agent 反复执行 stage and commit 这类操作烧 token,不如让它构建一个工具,后续交互全部免费;文中给出 Connect 4、Winget 包管理、SQLite 浏览器、Jekyll 编辑器等实例,并展示如何用 canvas 把 research→prototype→plan→implement→iterate→finalize 工作流中的自己移出循环。
来源:github.blog
Runway WorldPrompt 与实时世界模型的工程难题 | 与 Genie 3 / World Labs 的横向对比
Latent Space 独家采访 Runway CTO Kamil Sindi 与首席研究科学家 Robin Kahlow,解读 GWM Worlds 2 的新特性 WorldPrompt——一种指定生成世界及其内部动作的输入格式,可固定首帧等环境要素、生成带时间戳事件、实时 prompt 动作,相当于角色/相机/环境的控制层(类游戏 NPC 逻辑,但只是 prompting 而非脚本)。文章给出 Runway、Google Genie 3、Odyssey-2 Pro、World Labs RTFM 的横向对比表,并拆解把视频模型变成实时运行时的两大工程难题:从整段生成改为逐帧生成、以及把生成压到实时帧率。对关注 LLM 之外前沿方向的从业者,这是一手技术路线图。
来源:latent.space
SemiAnalysis ClusterMAX 3.0:GPU 云评级大更新,Nebius 与 CoreWeave 并列 Platinum | 323 家供应商、77 家 neocloud 深评
SemiAnalysis 发布 ClusterMAX 3.0,覆盖 323 家供应商(上版 209 家)、深度评测 77 家 neocloud、访谈 200+ 终端用户,按计算/网络/存储/编排/UI/监控/支持/安全等 10 大类打分。核心变动:Nebius 与 CoreWeave 并列 Platinum;Google Cloud 升入 Gold,Azure 降至 Silver,Crusoe 跌至 Bronze,Fluidstack 变为 Unavailable;新增「参与奖」层级容纳 15 家仅达最低标准的厂商,全球仅 19 家获 Medallion 评级。报告还拆解融资、Blackwell/Grace-Blackwell 部署、向 Vera Rubin 迁移、Scale-Out 网络、可靠性、安全与 Agentic Coding 趋势,附 3 万字逐家点评——选云、谈判与 TCO 建模的一手参考。
白宫要求 OpenAI/Anthropic 新模型先给美国测、再给英国 AISI | 前沿模型发布前的地缘政治准入博弈
POLITICO 报道白宫要求 OpenAI 和 Anthropic 在完成美国政府测试前,不得将新模型提供给英国 AI 安全研究所(AISI)。Anthropic 已同意,其 Claude Mythos 5.1 仅向美国部分伙伴开放,未给英国 AISI;英国 AISI 主任在致议会委员会的信中承认无法访问 Anthropic 模型,但强调仍对部分前沿模型有预发布访问权,并称已提前测试 OpenAI 的 GPT-6 Astra。英国呼吁建立全球共享原则,白宫则寻求对新模型的首测权——这条揭示了模型发布节奏、监管协调与英美 AI 安全合作裂痕,对理解「谁先拿到前沿模型」有直接价值。
来源:politico.com
包管理器沙箱与 coding agent 沙箱正在合流 | Homebrew 7.0.0 与 Claude Code/Codex 用同一套内核原语
Homebrew 7.0.0 把安装拆成有网络的 fetch 阶段与离线的 install 阶段,用内核级 Landlock 取代 Bubblewrap,并把任意 Ruby 的 post_install 迁移为签名声明式 *_steps。文章指出这套沙箱原语(Seatbelt/Landlock/seccomp)正是 Claude Code、Codex CLI、Cursor、Gemini CLI 采用的同一套栈,因为两者都在跑未审查代码。更关键的是二者已交叉:agent 会替你装包,Shai-Hulud 蠕虫通过 postinstall 脚本写入 Claude Code 配置实现持久化,Mandiant 报告攻击者劫持编码助手会话在约 100 个内部仓库扩散。作者给出 gating/confining/replacing 三分法,共同未解难题是验证受限阶段的输出——建立 agent 安全执行 mental model 的扎实素材。
来源:nesbitt.io
LiquidAI 发布 LFM2.5-VL-DSpark:视觉语言模型投机解码加速最高 3.13x | 仅加 8.9% 参数、day-one 支持 llama.cpp/MLX-VLM/SGLang
LiquidAI 为 LFM2.5-VL-3B 发布实验性 DSpark 投机解码 draft 模型:仅增加 280M 参数(+8.9%),在 M5 Max 上解码加速 2.30-3.13x、H100 上最高 2.66x,端到端最高 2.62x,且不改变输出质量。draft 模型为 4 层 attention-only 结构、block size 9,通过固定层隐藏态条件化生成候选 token,视觉 patch 与文本 token 投影到同一表征空间,因此推理算法与文本版一致。附六类视觉任务(VQA/图表/多轮对话)实测与视觉负载下投机解码的局限分析——对做 VLM 推理优化的团队是一份可复用的加速配方。
llmfit:按真实硬件给 5,372 个本地模型打分排序 | 拆解「参数量×字节数 vs VRAM」经验公式的四个失效点
llmfit 是一款本地 LLM 选型工具,读取真实硬件后对 5,372 个模型按 fit/speed/quality/context 四维打分并加权排序(chat 侧重速度、reasoning 侧重质量)。文章指出传统 VRAM 估算公式的四个失效点:KV cache 随上下文增长(Qwen3.5-9B 从 4K 的 6.6GB 涨到 131K 的 22.1GB)、MoE 稀疏激活(Mixtral 8x7B 实际只需 7.7GB 活跃权重)、混合硬件无法合并显存、以及「装得下但慢到不可用」。作者在 RTX 4080 上实测 68.2 tok/s,比预测的 40.9 高出 40%,并可用该实测值反向校正全机估算——本地部署/自托管团队可直接用其命令流(doctor/fit/recommend/info/bench)做选型。

🎙️ Podcast Picks

Who Feeds the GPUs? Inside AI's Hidden $30B Layer | Renen Hallak, VAST Data

📍 Source: The MAD Podcast | ⭐ ⭐⭐⭐⭐/5 | 🏷️ Infra, Agent, Interview | ⏱️ 01:10:14
VAST Data founder Renen Hallak dives into the data layer everyone ignores: what an AI factory really is, the DASE shared architecture, why traditional databases collapse at trillion-vector scale, and how training infra differs from inference. Also covers KV cache, RAG and agent memory, identity/permission security for agents, and DataEnclave confidential AI. Plus demand signals (customers going from 500PB to 2EB), circular financing risk, the neocloud landscape, and lessons from working with xAI and NVIDIA.
💡 Why Listen: If you think AI infra is just GPUs, this is the missing half. A $30B company CEO explaining why storage is the real bottleneck — and the agent memory angle is directly useful.

Runway's WorldPrompt and the Engineering of Real-Time Worlds

📍 Source: Latent Space | ⭐ ⭐⭐⭐⭐/5 | 🏷️ MultiModal, Research, Interview | ⏱️ 1:36:21
Runway launched GWM Worlds 2 and WorldPrompt, an input format for specifying generated worlds and their internal actions using autoregressive diffusion. CTO Kamil Sindi and chief research scientist Robin Kahlow explain how WorldPrompt controls characters, cameras, and environments like a game engine, and compare it against Genie 3, Odyssey-2 Pro, and World Labs RTFM. The core value is the engineering reality check: latency, continuous interaction length, and the limits of prompting as a control paradigm.
💡 Why Listen: Great if you care about multimodal generation or agent environments. The competitor comparison table alone is worth it, though it's less directly relevant if you're purely LLM-focused.

From AGENTS.md to Enterprise Deployment

📍 Source: Practical AI | ⭐ ⭐⭐⭐⭐/5 | 🏷️ Agent, Infra, Product | ⏱️ 48:52
How AI agents move from prototype to enterprise deployment: security, compliance, scalability, and reliability. Nick Kuhn shares VMware Tanzu's hands-on experience with agent build packs, MCP gateways, shared memory, identity, and sandbox isolation, plus lessons from years of platform engineering.
💡 Why Listen: Practical and grounded. If you're shipping agents inside a company, the MCP gateway and sandboxing bits are exactly the stuff that bites you in production.

AI Agents Are Moving Into the Real World

📍 Source: AI Daily Brief | ⭐ ⭐⭐/5 | 🏷️ Agent, Product, Regulation | ⏱️ 00:28:07
This episode focuses on agents entering real-world scenarios: Meta bringing Muse to smart glasses and new wearables, GrokBot turning Teslas into voice assistants, and whether consumers actually need agents. Headlines also cover Claude's early biology findings, Trump's "superintelligence" rebranding, and the UN AI governance debate.
💡 Why Listen: A quick 28-minute catch-up. Good for scanning the week's agent product news, but don't expect deep analysis.

Re-Founding Incumbents for the AI Era with Sequence Holdings Co-Founder and CEO Michael Lee

📍 Source: No Priors | ⭐ ⭐⭐/5 | 🏷️ Product, Funding, Interview | ⏱️ 42:44
Michael Lee talks with Sarah Guo about using AI to transform traditional industry leaders from the inside rather than disrupting them. His thesis: traditional consulting and software sales models can't deliver enterprise AI transformation — you need a holding company structure plus deep engineering integration. He shares results from the $7.7B take-private of Baldwin and the BankSouth investment, and why permanent holding structures favor long-term compounding.
💡 Why Listen: An operator's view on AI in legacy industries. Concrete cases, but it's more investing and ops than LLM tech.

📄 Paper Highlights

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

Texas A&M University, AWS AI, Amazon | 🏷️ Agent Framework, Fine-tuning, RLHF/DPO
Shows on-policy self-distillation makes multi-turn agents act confident without the underlying knowledge — sometimes worse than the base model. Their fix moves privileged info from the loss into the sampler.

Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

Alibaba | 🏷️ Agent Framework, Agentic Workflow, Tool Use
Alibaba's closed-loop AI-for-AI pipeline links data production, training, and deployment for mobile planner agents, with a reward-engineering trick to cut reasoning and tool-use costs.

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

Nubank | 🏷️ Agent Deployment, Agentic Workflow, Safety
Nubank screened CX agents in simulation before shipping, then hit a 36.69-point tNPS lift in live A/B tests — a rare look at agentic deployment with real production numbers.

🐙 GitHub Trending

Fuzzing Taskflow | LLM-driven fuzzing with clean separation
GitHub Security Lab's reference implementation lets an LLM agent run the whole C/C++ fuzzing loop: find entry points, write harnesses, run AFL++, read coverage, triage crashes. The agent only decides; MCP tools expose primitives like run_afl_for, and state lives in SQLite. A ready-to-run template for agentic pipelines.
GitHub | ⭐ New | 🗣️ Python | 🏷️ Security, Agent, Fuzzing
GameDev OS | Nine experts, 64 skills for game agents
Scenario open-sourced a coding-agent setup with nine specialist roles — 2D/3D art, environment, audio, video, marketing, tech director, cost manager — each wired to models like Seedance, GPT, Meshy, and Rodin. Chains from concept art to modeling to scene to audio to trailer.
GitHub | ⭐ New | 🗣️ Multi | 🏷️ Agent, GameDev, Multi-Agent
Quail | Open-source AI-SQL engine
Built with Modal, Quail co-schedules query planning and LLM inference, hitting 1B+ input tokens per minute on a single H100. Targets the new inference workload of AI data operators.
GitHub | ⭐ New | 🗣️ Rust | 🏷️ Database, Inference, SQL
  • AI
  • Daily
  • Tech Trends
  • AI Tech Daily - 2026-09-26AI Tech Daily - 2026-09-24
    Loading...