AI Tech Daily - 2026-09-22
2026-9-22
| 2026-9-22
字数 3148阅读时长≈ 8 分钟
type
Post
status
Published
date
Sep 22, 2026 05:00
slug
ai-daily-en-2026-09-22
summary
Xiaomi open-sourced MiMo-V2.6 Pro and Flash, a 1.02T-parameter multimodal family with a 1M context window and the highest AA Intelligence Index of any open model at 46 — plus the RL stack, environments, and distilled Qwen3.5-9B weights. StepFun's Step 5 Preview matched Kimi K3 at 44 on the same inde
tags
AI
Daily
Tech Trends
category
AI Tech Report
icon
📰
password
priority
1

📊 Today's Overview

Xiaomi open-sourced MiMo-V2.6 Pro and Flash, a 1.02T-parameter multimodal family with a 1M context window and the highest AA Intelligence Index of any open model at 46 — plus the RL stack, environments, and distilled Qwen3.5-9B weights. StepFun's Step 5 Preview matched Kimi K3 at 44 on the same index for roughly 2.8x less per task. Meanwhile OpenAI disclosed its internal model solved 100+ open math problems and stood up an independent advisory group, AGMAI, to review results.

🔥 Trend Insights

  • Open-weight race goes multimodal: Xiaomi's MiMo-V2.6 Pro/Flash eat text, images, video, and audio in one 1M-context model, and ship with the full RL stack — open labs are now releasing training pipelines, not just weights.
  • Cost per task becomes the metric: Step 5 Preview ties Kimi K3 on intelligence at ~2.8x lower cost, while Grok 4.7 loses its price edge to Astra — buyers are now comparing dollars per task, not just scores.
  • Decision models as infrastructure: LangChain and Bespoke Labs both stress-tested TypeSafe AI's Jev within a day, showing non-generative "System One" models may become a third path alongside LLM-as-judge and code-based evals.

🐦 X/Twitter Highlights

📈 热点与趋势

  • 小米开源 MiMo-V2.6 Pro / Flash,同批放出 RL 栈 - Pro 为 1.02T 总参 / 42B 激活,Flash 为 309B / 15B 激活,单权重吃文本、图像、视频、音频,1M 上下文,AA 智能指数 46 分为开源最高。同批开源用 MiMo 轨迹蒸馏的 Qwen3.5-9B、7K 环境和 RL 训练代码。团队称这是按算力计开源团队做过的最大单次 RL 训练之一,数十人投入 @XiaomiMiMo @_LuoFuli(小米 MiMo 团队) @teortaxesTex(AI 内容博主)
  • Step 5 Preview 拿 44 分、每任务 $0.71 - Artificial Analysis 评测:AA 智能指数 44,与 Kimi K3 (max) 同分,但每任务成本约低 2.8 倍;HLE 46%、CritPt 21%。Agentic 侧落后,GDPval-AA 1,566 Elo,低于 Qwen3.8 Max 的 1,668。定价 $1/$2.70 每 1M tokens,权重 10 月 15 日开源 @StepFun_ai @ArtificialAnlys(模型评测机构)
  • OpenAI 内部 AI 已接手实验模型训练 - The Information 报道:内部模型能写 GPU kernel、优化训练代码,研究者给一个优化示例后让 AI 自己跑数周;过去要数年的实验现在约一周做完,内部 agent 之间直接协作、不经人类;另有说法称 24 天内解决 100 多个数学开放问题 @AISafetyMemes(AI 安全话题账号,转述 The Information)
  • Boston Dynamics 与 HMG 建机器人应用中心 - 在 HMG Metaplant America 设 Robotics Metaplant Application Center(RMAC),作为把 Atlas 机器人接入汽车制造产线的测试与训练基地 @BostonDynamics(机器人公司)
  • Jeff Dean 离职后首次公开对话 - Dawn Song(UC Berkeley 教授)主持,谈从 MapReduce、TensorFlow、TPU 到 Gemini 的路径选择,RSI 可能的形态、科研发现循环自动化、自主 AI 安全,以及 Jeff Dean(Google 前首席科学家,任职 27 年)新创业公司的方向 @JeffDean @dawnsongtweets
  • Grok 4.7 每任务成本反超 Astra - 在 Artificial Analysis 上,Grok 4.7 不再便宜,每任务成本高于 Astra @theo(Theo,Web 开发者 / 播客主)

🔧 工具与产品

  • Halo:后训练吞吐达 TRL 的 2.8 倍 - White Circle 开源 Halo,峰值内存更低,模型保持原生 HuggingFace 格式,可直接接管开源模型后训练 @whitecircle(社区开发者)
  • Husky 推理引擎称比 Apple MLX 快 4.5 倍 - Underdog 的 Model-Specific Inference 引擎在 MacBook 上跑到 730 tokens/s,主打本地私有模型 @0xSigil(Underdog 团队成员)
  • Kimi K3 上线 Amazon Bedrock - 可用 Bedrock 的加密与审计控制跑编码、文档分析和长时 agent 工作流,支持显式提示缓存 @Kimi_Moonshot(月之暗面)

⚙️ 技术实践

  • SGLang 在 Blackwell 上用 NVFP4 KV cache - 每 token KV 占用降到 FP8 的约 56%,上下文容量 1.78 倍,32K / 160K / 1M 解码吞吐分别 +37% / +58% / +78%;GPQA-Diamond 与 AIME 2025 近无损。方案含两级缩放、解码期 kernel 内反量化与分页 KV,一个 flag 开启 @lmsysorg(SGLang 开源推理引擎出品方)
  • vLLM day-0 支持 MiMo-V2.6 双尺寸 - 沿用 MiMo 的混合滑窗 + 全局注意力、推理与工具调用 parser、DFlash 投机解码(每步起草 7 token),原生 FP8 权重 @vllm_project(开源推理引擎)
  • vLLM 调优 Qwen3.8-2.4T 的 Pareto 前沿 - 在 GB300 NVL72 上,8K 输入 / 1K 输出下高吞吐侧 5K tokens/s/GPU、低延迟侧 180 tokens/s/user;流程为预算 KV cache、prefill 与 decode 分开压测,再为每个服务目标选拓扑和 MTP 设置,部署配置已放出 @vllm_project(开源推理引擎)
  • 测试期通信:Team-of-N 何时强过 Best-of-N - Dimitris Papailiopoulos(ML 系统研究者,威斯康星大学麦迪逊分校教授)等让 N 个相同 agent 无分工、只靠共享文本日志协作:ARC-AGI-3 上 Team-of-5 Sonnet-4.6 追平 Best-of-33,其中一个 64 次单 agent 都没解出的游戏被解出 65%;polyomino packing 上 Team-of-3 Opus 4.6 超过 best-of-60;MNIST 压缩上四人 GPT-5.6 Sol 队找到 1,957 字节、99.4% 准确率的模型,比最好人类解小 20% @DimitrisPapail

⭐ Featured Content

Jev ecosystem takes shape in a day: from "decision model" to agent evaluator, third-party tests and open-source baselines land together | Non-generative models start getting validated as infrastructure
TypeSafe AI's Jev — a "System One" model that only outputs calibrated decisions and generates no text — drew dense third-party validation within a day of launch. LangChain tested it as an agent evaluator: scoring variance was 92–913x lower than GPT-5.6 Luna/Terra and Claude Sonnet 4.6, averaging 0.44s and $0.00035 per call ($0.34 total vs Claude's $28.17). The author stresses the test scope was narrow and conclusions premature, but notes the "decision-first" architecture could become a third path alongside LLM-as-judge and code-based eval. Meanwhile Bespoke Labs open-sourced a comparable model, Bespoke-Nimble-9B (Qwen3.5-9B LoRA, Apache-2.0), reading logits directly for enum/boolean answers instead of generating JSON, and ran the first head-to-head against human labels — Jev won by just 1.2 macro points, with latency of 106ms on H100 / 444ms on M5 Pro vs Jev API's 247ms. They also released data, training recipe, and a human-labeled benchmark. For teams doing agent routing, LLM-as-Judge, or eval cost optimization, this is the key day the "decision model" track moved from single launch to reproducible comparison.
Stratechery reverse-engineers Anthropic's "Pacing the Frontier": three overhangs behind the safety narrative | A first-hand strategic breakdown of frontier lab business logic
Stratechery dissects Anthropic's slowdown argument from a strategy rather than philosophy angle: on the surface it's safety concern, but in substance it eases several overhangs frontier labs face — capability overhang from models improving too fast, the agent paradigm's extreme hunger for tokens, and the differentiation moat from model-plus-harness integration. Drawing on his own hands-on agentic coding experience, including attempts to build his own harness, the author argues the "model commoditization" bet may fail and profits will flow to companies doing model+harness integration. He cites a Nadella interview as evidence of the temporary nature of Microsoft's E7 and Claude Cowork partnership. Good for practitioners tracking the agent industry landscape and frontier lab business logic.
OpenAI discloses internal model solved 100+ open math problems, forms independent advisory group AGMAI | Designing an academic buffer mechanism after capability exceeds expectations
OpenAI disclosed that its internal model, which began training on August 28, has solved 100+ open math problems beyond the Navier–Stokes Millennium Prize problem — progress so fast it surprised internal mathematicians. In response, OpenAI and mathematicians co-founded an independent advisory group, AGMAI (affiliated with IAS), with members including Timothy Gowers, Martin Hairer, Edward Witten, and Ravi Vakil. Key design: advisors take no OpenAI salary, can voice opinions proactively, can publicly criticize OpenAI, and explicitly are not tasked with "advising OpenAI to slow down internal math progress" — only result review, dissemination coordination, and academic standards advice. A first-hand sample of how AI companies build buffer mechanisms with academia when capability exceeds expectations.
Sources: openai.com
Nathan Lambert's congressional briefing goes public: Chinese open-weight models lead by roughly 2-5 months | Quantifying the US-China open-source gap with two hard datasets
The Interconnects author turned a briefing to US congressional members and staff on open-weight models into a public article, systematically laying out the lineage differences among open-weight, open-source, and closed models. He uses two datasets — Hugging Face downloads (China's cumulative 3.2B, roughly twice the US) and the Artificial Analysis Intelligence Index (GLM-5.3 at 45, Kimi K3 at 44, versus the strongest US open models Thinking Machines Inkling at 26 and Nemotron 3 Ultra at 23) — to show Chinese open-weight models lead by roughly 2-5 months. The piece also names the genuinely open US camp (AI2 Olmo, OpenAthena Marin, EleutherAI Pythia) and Nvidia Nemotron's semi-open positioning. A first-hand framework for understanding the US-China open-source competition.
SemiAnalysis breaks down MoE inference data flow: four stages with sharply different compute intensity | A system-level map for inference serving capacity planning
SemiAnalysis's long piece splits the data flow of mapping MoE models onto inference hardware into four stages: prefill, midfill, decode attention, and decode experts. It notes the four have sharply different demands on compute, memory, and network bandwidth — prefill is compute-intensive, while decode is low in compute intensity but dominated by weight movement. The article stresses that treating the four stages as one workload throws away MoE's structural advantages, and discusses the tradeoffs of aggregated vs disaggregated serving, touching on orchestration layers like Dynamo, Mooncake, and vLLM. A rare system-level map for teams doing inference optimization, capacity planning, and expert parallelism.
Tim Dettmers announces dlab open-source week: 550B model runs on a 128GB MacBook | Must-follow for local inference and agent harness this week
Tim Dettmers announced dlab open-source week, arguing "the unit of research is no longer the paper but the ecosystem": in the agent era a single paper becomes cheap, and the hard part is shipping a coherent ecosystem of mutually reinforcing releases. He previews three things — frontier autonomous research, the most efficient test-time scaling, and auto-compaction more efficient than Claude Code/Codex. Hard numbers already given: after agents autonomously optimized Metal kernels, Qwen 3.6 35B-A3B runs at 450 tok/s at 1.5 bits/weight; Qwen 3.8 Flash Next 125B runs on a single 24GB GPU, and DeepSeek V4.1 550B runs on a 128GB MacBook. A must-follow this week for anyone tracking local inference, agent harness, and quantized deployment.
Benchling publishes multi-tenant agent code execution defense-in-depth: DNS egress is the covert channel | A directly reusable isolation architecture checklist
Benchling runs AI-agent-generated life sciences code on Amazon Bedrock AgentCore Code Interpreter, with 600+ execution sessions per day, 250+ tenants per week, and zero security incidents. Core insight: traditional sandboxes only block HTTP and outbound ports, but DNS resolution is often allowed by default and lacks visibility — exactly the covert channel for data exfiltration. The solution: a separate "Untrusted Code Account" hosting the ACCI VPC (no IGW/NAT), Route 53 Resolver DNS Firewall with three-tier policy (10 blocks malicious domains, 100 allows only whitelisted endpoints, 200 default deny), VPC Endpoint as the only network path, per-task STS injection of least-privilege credentials to avoid per-tenant IAM role explosion, and a continuous validation suite simulating exfiltration attacks. A directly reusable architecture checklist for teams doing multi-tenant agent execution isolation.
MCP 2026-07-28 revision explained: Streamable HTTP fully removes protocol-level sessions | A migration decision table for a breaking change
The 2026-07-28 MCP revision is a breaking change: the Streamable HTTP transport fully removes protocol-level sessions and Mcp-Session-Id, drops the initialize handshake, HTTP GET stream, and stream resumption, and moves to per-request self-description — protocol version, client identity, and capabilities are passed per-request via _meta, and application state across tool calls uses explicit handles as tool arguments. The article goes item by item through "can delete / must add / keep for compatibility," gives three implementation generations (legacy 2025-11-25, modern, dual-era), and includes an integration decision table and production checklist. It also flags where SDKs lag the spec (TS SDK split into the @modelcontextprotocol/* 2.0.0 family, Python 2.2.0). It notes the governing body is the Agentic AI Foundation, with a new lifecycle policy of a minimum 12-month deprecation window.
Sources: dev.to

🎙️ Podcast Picks

Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI

📍 Source: Latent Space | ⭐ ⭐⭐⭐⭐⭐ | 🏷️ LLM, Agent, Interview | ⏱️ 2:20:53
Diogo Almeida (InstructGPT co-author, TypeSafe AI CEO) lays out the logic behind Jev's creation: he argues that after ChatGPT's success, the industry was left with only one path — autoregressive chat fine-tuning — which skewed alignment, refusal, and reliability. He critiques all three types of RLHF as the wrong north star, and advocates for small production-facing System One models, stressing that KV cache decides everything and models should fade into the background like regex. The episode also covers Jev's deployments in coding agents, voice + browser control, entity resolution, and natural language search, and previews the ReasoningJev direction.
💡 Why Listen: An InstructGPT co-author explaining why he thinks the whole industry took a wrong turn. Sharp, opinionated, and full of concrete design tradeoffs — worth it for anyone building agents or doing LLM engineering.

Why Scaling Prediction Cannot Create Intelligence - Alexander Mattick

📍 Source: ML Street Talk | ⭐ ⭐⭐⭐⭐ | 🏷️ Research, LLM, Agent | ⏱️ 02:14:20
Alexander Mattick connects modern machine learning through the lens of inference, covering Monte Carlo, GFlowNets, energy models, diffusion, and flow matching. He's blunt: energy model sampling is poor value for money, and JEPA and world models are closer to brands than technical categories. The back half digs into deep learning theory, constrained RL, cybernetics vs RL, the Bitter Lesson, and the nature of world models, arguing "prediction is not control."
💡 Why Listen: Hardcore theory, but the kind that reshuffles how you think about generative models. Great if you want a mental model beyond scaling.

The State of the AI Debate

📍 Source: AI Daily Brief | ⭐ ⭐⭐ | 🏷️ Regulation, Funding, Research | ⏱️ 00:27:12
This episode covers the AI policy debate: Trump's proposal for an AI Force, political pushes for kill switch legislation, calls for AI slowdown, and the China factor. It focuses on the outlook for US-China AI cooperation ahead of a Trump-Xi meeting. Headlines include Anthropic's IPO delay, new bio labs, and signs of stress in the data center debt market.
💡 Why Listen: A quick 27-minute policy roundup. Light on tech, but useful for tracking regulation and funding signals.

📄 Paper Highlights

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

Meta, CMU | 🏷️ Agent Deployment, Evaluation, Inference
First-hand deployment report on recurring evaluation for a production analytics agent serving tens of thousands of monthly users — and why the team shipped simpler fixed subsets over theoretically better adaptive testing.

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Xiaomi, PKU, HKU, Renmin University of China | 🏷️ Code Agent, RL, Training
Turns source code alone into executable RL environments — no issues or commits needed — yielding 5,545 tasks across 3,185 repos and lifting MiMo-V2.5 on five coding benchmarks.

NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

NVIDIA | 🏷️ Multimodal, Tool Use, Architecture
An open full-duplex speech model that listens, transcribes, reasons, calls tools, and speaks in one streaming architecture — a system-level first for combining real-time conversation with native tool use.

🐙 GitHub Trending

Halo | Post-training at 2.8x TRL throughput
White Circle's open-source post-training framework hits 2.8x TRL's throughput with lower peak memory, and keeps models in native HuggingFace format so you can take over open models directly. A drop-in speedup for anyone fine-tuning.
GitHub | ⭐ 3,100 | 🗣️ Python | 🏷️ Training, Fine-tuning, LLM
Husky | Model-specific inference engine for Macs
Underdog's Model-Specific Inference engine claims 4.5x faster than Apple MLX, hitting 730 tokens/s on a MacBook for local private models. Worth a look if you run models on-device.
GitHub | ⭐ 2,400 | 🗣️ Rust | 🏷️ Inference, Local, Mac
Bespoke-Nimble-9B | Logit-reading decision model baseline
Bespoke Labs' open baseline (Qwen3.5-9B LoRA, Apache-2.0) reads logits directly for enum/boolean answers instead of generating JSON, with data, training recipe, and a human-labeled benchmark released. The first reproducible comparison point for the "decision model" track.
GitHub | ⭐ 1,800 | 🗣️ Python | 🏷️ LLM, Evaluation, Agent
  • AI
  • Daily
  • Tech Trends
  • AI Tech Daily - 2026-09-23AI Tech Daily - 2026-09-21
    Loading...