AI Weekly 2026-W41
2026-10-10
| 2026-10-10
字数 5365阅读时长≈ 14 分钟
type
Post
status
Published
date
Oct 10, 2026 05:41
slug
ai-weekly-2026-W41-en
summary
One event dominated this week: on October 6, OpenAI published 722 mathematical manuscripts on GitHub. The results came from an undisclosed internal frontier model, spanning algebra, number theory, topology, theoretical computer science, and mathematical logic. Out of 4,000 evaluated problems, it produced results for 90 of the Top 500 open problems across all of mathematics — at an average compute cost of roughly three hours of ChatGPT Pro per result. Princeton's IAS AGMAI issued a formal statement. Anthropic's Levent Alpöge called Result 003, "Quasi-Riemann Hypothesis," "the most important moment in the history of mathematics." At the same time, three retractions, roughly 20% of results being counterexamples, and disputes over Lean verification are all part of the story. If there's one thing worth watching this week, it's not "AI can do math" — it's that "can AI-produced math be verified" became a public argument for the first time. The second thread is how densely the inference stack is adapting to new hardware. NVIDIA Vera Rubin picked up Day 0 support from both vLLM and SGLang within days: vLLM reported MiniMax M3 hitting 7.8x the throughput of GB200 on AgentX, while SGLang reported FP8 MLA speeding up 20% at batch 1 / 128K context and MoE tail fusion cutting 276 kernel launches per step. In the same week, vLLM shipped v0.31.0 (717 commits, 307 contributors) and released the vLLM-Omni technical report for its omni-modal serving runtime. Hardware generations and engine iterations are now meshing on a weekly cadence. The third thread is agent engineering shifting from "can it run" to "can it be governed." GitHub disclosed the impact of agent-scale development on Git infrastructure (monthly commits up 5x year-over-year, pushes up 4.9x). Microsoft turned agent sandboxing into an OS-level Windows primitive (MXC GA). Anthropic launched Claude Managed Agents dynamic workflows in public beta. And several papers directly study agent time budgets, judge reliability, and si
tags
AI
周报
category
AI Tech Report
icon
password
priority
1

📊 Weekly Overview

One event dominated this week: on October 6, OpenAI published 722 mathematical manuscripts on GitHub. The results came from an undisclosed internal frontier model, spanning algebra, number theory, topology, theoretical computer science, and mathematical logic. Out of 4,000 evaluated problems, it produced results for 90 of the Top 500 open problems across all of mathematics — at an average compute cost of roughly three hours of ChatGPT Pro per result. Princeton's IAS AGMAI issued a formal statement. Anthropic's Levent Alpöge called Result 003, "Quasi-Riemann Hypothesis," "the most important moment in the history of mathematics." At the same time, three retractions, roughly 20% of results being counterexamples, and disputes over Lean verification are all part of the story. If there's one thing worth watching this week, it's not "AI can do math" — it's that "can AI-produced math be verified" became a public argument for the first time.
The second thread is how densely the inference stack is adapting to new hardware. NVIDIA Vera Rubin picked up Day 0 support from both vLLM and SGLang within days: vLLM reported MiniMax M3 hitting 7.8x the throughput of GB200 on AgentX, while SGLang reported FP8 MLA speeding up 20% at batch 1 / 128K context and MoE tail fusion cutting 276 kernel launches per step. In the same week, vLLM shipped v0.31.0 (717 commits, 307 contributors) and released the vLLM-Omni technical report for its omni-modal serving runtime. Hardware generations and engine iterations are now meshing on a weekly cadence.
The third thread is agent engineering shifting from "can it run" to "can it be governed." GitHub disclosed the impact of agent-scale development on Git infrastructure (monthly commits up 5x year-over-year, pushes up 4.9x). Microsoft turned agent sandboxing into an OS-level Windows primitive (MXC GA). Anthropic launched Claude Managed Agents dynamic workflows in public beta. And several papers directly study agent time budgets, judge reliability, and simulation training environments. The question that kept surfacing this week: once agents run at scale, who sets their boundaries.

OpenAI's Math Papers and AI Scientific Discovery

This has to come first this week, because its impact extends well beyond mathematics.
AINews' roundup (Latent Space) gives the most complete first-hand account: 722 manuscripts, 372 result families, from an evaluation of roughly 4,000 research problems, averaging about 3 hours of ChatGPT Pro compute per result — compared to 88 hours / 10,000 agents for the earlier Navier-Stokes batch. That magnitude difference alone is worth noting: compute efficiency improved by more than an order of magnitude. Specific results cited include integer multiplication faster than n log n, uniqueness for elastic inverse problems (a 3D open problem dating to 1994), partial progress on Riemann/Hodge/BSD, and an explicit note that roughly 20% of results are counterexamples — a key point, since counterexamples are also mathematical output but easily lost in the retelling.
Zvi's deep follow-up offers a calmer, item-by-item breakdown. He categorizes the 719 papers (after three retractions) by result type, interprets the significance of each — tightened Riemann bounds, matrix multiplication, integer multiplication, Pi, Unique Games, Hodge/Birch, Hilbert's tenth problem, Hadwiger graph coloring — and records three things: the shock to the math community, the Lean verification controversy, and the three retractions. The retractions haven't been sufficiently emphasized in mainstream coverage. In a release claiming 722 results, three retractions mean the verification mechanism is working — and also that it doesn't always catch errors on the first pass. Zvi's conclusion is more conservative than AINews': the results are interesting, but "whether they're useful" is a separate question.
Background. New Scientist ran two pieces within two days — one on the scale of the release, one picking out "the most interesting findings." The latter notes the results span "both long-standing mathematical puzzles and tiny improvements to niche problems rarely considered" — two categories that are indistinguishable in a Turing-test-style evaluation but differ enormously in how mathematicians value them. DataCamp's writeup confirms the October 6 release date and that the model's output cycle was "measured in weeks." AGMAI's statement is the most quotable passage: the impact "extends beyond individual results to the entire field of mathematics and the mathematical community."
Looking at the broader AI for Science picture, the Periodic Labs interview lays out a complete case for a different path. Liam Fedus (ChatGPT co-creator) and Ekin Dogus Cubuk (author of DeepMind's GNoME/MatterGen) argue that "intelligence alone is insufficient for science" — scientific discovery differs fundamentally from math and programming, requiring reasoning under noise, uncertainty, and missing information, with experiments as the ultimate ground truth. Their proposed synthesis superintelligence is a heavier approach: move RL environments into the physical world, use AI for materials characterization, DFT simulation, and high-throughput autonomous labs, and train models to learn "the process of doing science" rather than just studying published results. One judgment deserves to be pulled out on its own — failed experiments and negative results may be the most valuable training data. The contrast with this week's math results is striking: in math, counterexamples are output; in science, negative results are training data. Both require systems that can accommodate "unsuccessful" information — precisely what most current benchmarks and reward designs handle poorly.
Put the two together, and this week offered two non-intersecting paths for AI for Science: dense search within pure symbolic space (OpenAI's math), and a closed loop that must land in physical experiments (Periodic Labs' materials). The former has already delivered 722 papers' worth of volume; the latter is still building infrastructure — and both founders repeatedly stress that future frontier models will still have to run physical experiments. Which path gets there first is too early to call.

GPU Cluster Scheduling and Next-Gen Inference Stack Adaptation

The keyword for the inference stack this week is "Vera Rubin ready," but the more durable value comes from an Ai2 engineering retrospective.
Ai2's GPU cluster scheduling retrospective (Hugging Face Blog) describes how they turned "who gets how many GPUs" from case-by-case operational negotiation into a transparent administrative budget process. Start with the pathologies — this part is more valuable than the solution itself: GPU squatting (users holding no-op workloads while waiting to debug), priority inflation (100% of workloads marked HIGH, starving low-priority jobs), and on-call engineers spending large amounts of time negotiating shutdowns of non-preemptible workloads. These three will ring true on virtually any self-managed cluster. The solution is GPU time budgets (per-project time allocations) + hierarchical fair-share + time-slicing contracts, with an evaluation framework built as a four-tier metrics pyramid: availability → occupancy → impact → utilization. The paper validates with both simulation and real cluster results. This is rare first-hand material that includes failure modes and a migration path — most cluster scheduling writeups only describe the final architecture, never why the old one broke.
Hardware side. vLLM supports Vera Rubin (vLLM Project) reports MiniMax M3 hitting 7.8x the throughput of GB200 on AgentX, because Rubin is compatible with most Blackwell kernels — DeepSeek, Kimi, MiniMax, GLM and others are Day 0 ready. Author Woosuk Kwon adds that beyond compatibility, a batch of Rubin-specific optimizations also shipped. SGLang's corresponding adaptation provides finer kernel-level data: FP8 MLA speeds up 20% at batch 1 / 128K context, KDA verification speeds up 20% with bitwise-identical output, and MoE tail fusion cuts 276 kernel launches per decode step for a 5.9% end-to-end speedup. The LMSYS official account also mentions Miles using SGLang for end-to-end RL rollout on Rubin, including 64 concurrent sandboxes running on Vera CPUs.
On releases. vLLM v0.31.0 (717 commits, 307 contributors, 96 first-time contributors) highlights four areas: FlashMLA mega attention + NVFP4 KV cache + DeepGEMM sparse MQA for DeepSeek-V4.1-Flash; fast restart (vllm preload keeps weights in GPU memory across restarts); draft-model speculative decoding and DSpark adaptive verification in Model Runner V2; and MoonEP balanced all-to-all plus DeepEPv2 for large-scale serving. The vLLM-Omni technical report takes a different path — a unified serving runtime for omni-modal generation. Its motivation is stated plainly: voice assistants, visual generation, world models, and robot loops each have distinct execution patterns (multi-stage AR pipelines, iterative diffusion, stateful cross-step sessions), and LLM servers and diffusion stacks each excel at only one — leaving deployment to stitch together unrelated runtimes. vLLM-Omni uses an orchestrator to advance requests across stages, specialized engines to run computation, and connectors to pass payloads, letting duplex, world models, and robot loops share a single session path.
Two papers worth reading. RaReCache (USC / UC Irvine / Intel Labs) tackles precision decay in cross-model KV cache reuse. Prior work showed KV caches can be translated between same-family models via closed-form linear mappings, but transfer precision degrades as model size gaps widen. This paper's observation: transfer failures concentrate on a small number of information-dense tokens. The method uses rank disagreement as a scoring metric and selectively recomputes only critical positions — Qwen3-0.6B to 14B (23x parameter gap) needs only 30% of positions recomputed to retain 95-99% accuracy; Llama3-8B to 70B (8.8x) needs 40% recomputation to retain 96.5%. On the serving side, throughput reaches 1.8x that of direct prefill on the target model on a single GPU, and under saturated load a 30% recompute budget cuts median TTFT by 5.0x and P99 by 6.4x. The significance: small models prefill for large models, and large models only recompute critical tokens.
The other is a training-side numerical issue. Chen's FlashAttention-3 BF16 training anomaly: training an LLM with BF16 FlashAttention-3, training looks normal for a long stretch, then gradient norms spike 1000x, and final loss ends up 0.2 nats higher than FP32 attention — with no NaN anywhere in the run. The problem traces to the attention backward pass. This kind of "silent degradation" is harder to diagnose than a crash — no NaN means most health checks won't catch it. Anyone running long training jobs should spend time with that paper.
On the edge side, lithos-metal is open-sourced (Jia Zhihao), using megakernel + DSpark speculative decoding to push Qwen3.8-27B to a peak of 200+ tokens/s/user on a single Apple M5 Max. That throughput number on a single machine, single card, consumer hardware would have been a data center configuration a year ago.
Taken together, these point in a common direction: the optimization center of gravity in the inference stack is shifting down from "the model itself" to kernels, scheduling, and cache orchestration. Vera Rubin's Day 0 support rests on kernel compatibility. vLLM's major release focuses on KV cache formats and all-to-all balancing. RaReCache decides whether to compute at the token level. And Ai2's post defines the problem as an administrative one. The model weights haven't changed — every layer of the system around them is being rewritten.

Agent Evaluation, Time Budgets, and Simulation Training Environments

Several papers this week converge on one conclusion: agent failures are often not about capability, but about nobody telling the agent where the boundaries are — and the agent never learning how to use them.
On the Clock (AWS / UIC / NYU) directly studies agent behavior under explicit wall-clock time budgets. The setup: Qwen3.6-27B on five MLE-Bench Lite competitions, Qwen3-4B on Zork I. The first finding is deflating: declaring a time budget in the prompt alone doesn't let the agent translate that budget into controlled time usage. Three reasons — the harness provides no timing feedback, the agent can't reliably estimate action durations, and there's no learned mapping from "available time → appropriate strategy." Two classes of intervention: harness-level exposure of timing information with enforced deadlines, and RL with budget-aware rewards. Results diverge sharply: injecting timing information substantially improves Qwen3.6-27B's budget compliance with no performance loss; GRPO achieves near-perfect budget compliance on Zork I and generalizes to budgets unseen in training, but doesn't improve task performance on MLE-Bench. The most valuable conclusion is the last one: even when agents learn to respect budgets, they still don't use the extra time to perform better — the RL policy learns when to stop but often fills remaining time with repetitive actions, and multi-budget GRPO training collapses to the policy learned at the shortest budget. This gap between time compliance and effective time allocation is the core challenge for budget-conditioned agents.
AgentHorizon (ServiceNow Research / Mila / McGill and others) studies a different boundary problem: who decides whether the agent did it right. The benchmark contains 1,373 computer-use tasks drawn from 166 hours of human-recorded trajectories across three operating systems. One clever design choice: constructing negatives by swapping in similar instructions, testing whether a judge can distinguish "completed the task" from "completed a similar but incompatible request." Eleven judges were evaluated in two modes (passing full trajectories directly, or acting as a coding agent across five harnesses). Results: the best agentic judge (GPT-5.5) reaches 80.9% balanced accuracy on the AH subset; tool use helps some models but makes open-weight models perform worse; and judges vary enormously in their ability to accept valid trajectories versus reject failed ones. The problem framing is sharp: a trajectory with 300 screenshots and actions may look complete while actually violating instruction constraints or introducing side effects — and the judge needs to actively locate and verify evidence that is often hidden.
Who Verifies the Verifier? (AWS / HSBC) pushes this further — making the verifier itself the object of evolution. The approach expresses verifiers as checkable compositions of small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, with selection criteria based on agreement with ten anchored reference sets plus consensus on unlabeled outputs — explicitly not using agent scores as a selection criterion. On MBPP+, it improves held-out agreement by +0.21 over hand-written seed compositions, and holds on every seed. The finding most worth remembering is a warning: removing the anchoring guard causes the verifier to degenerate into an always-pass invalid scorer — and this degenerate verifier trains skills just as well. The implication: downstream task scores simply cannot certify a self-evolving verifier. The authors also report that Double Ratchet retains 88-110% of the lift from ground truth or rubrics across three settings: code generation, enterprise text-to-SQL, and reference-free report generation.
Two environment papers. StoreBench (AfterQuery) is a live business operations environment: the agent uses a production-grade backend of a mid-sized online clothing store, making decisions through the 29 tools a human operator would use, while customers order around the clock, suppliers change prices and go out of stock, and market shocks arrive with partial or zero warning. Three design details stand out: windowed action budgets make simulated time a function of action count (so model latency can't affect simulated time), pass thresholds are calibrated against scripted anchor policies, and rewards are hardened against a catalog of reward hacks. Seven frontier models run 11 scenarios; the best, DeepSeek-V4-Pro, passes 49% of task-seed cells, with human experts at 0.708 versus the best model at 0.700 — no model matches the scripted smart-triage heuristic (97%) on average. On GRPO post-training, Qwen3.5-27B raises held-out average composite score from 0.136 to 0.373 using only five disjoint tasks.
Synthesis Through Simulation (SAP Labs) takes the "generate training data" direction: LLM agents execute operations in a policy-enforced simulated enterprise environment to generate data. The core design is schema-free — because data is generated through the same environment that defines "what counts as legal," structural validity is guaranteed by construction, decoupling validity enforcement from distribution modeling so each can be handled independently. The Generalist Populator achieves 0.88 average marginal fidelity and 100% constraint satisfaction across ten environments — without accessing the database schema. By contrast, statistical synthesizers fail to apply on seven environments due to missing seed data, and agents with schema access fail 82% of trajectories on the airline environment. The framework, ten environments, and generated datasets are all open-sourced at https://github.com/SAP/synthesis-through-simulation.
On retrieval, RIT-RAG (IBM / IIT Kharagpur) addresses structure awareness in agentic RAG. Methods like PageIndex can navigate document structure but don't scale to large corpora — the full tree won't fit in an LLM context, so a document retriever must first pin down a single document, and a wrong pick is unrecoverable. RIT-RAG's approach: build a tree offline for each document (table of contents or sitemap), then at query time retrieve a batch of chunks and use their positions to induce a manageable subtree in reverse (possibly spanning multiple documents), after which an LLM agent navigates the subtree, reads selectively, and rewrites the query when needed. The division of labor is clean: retrieval decides where to look, the agent decides what to read. It achieves the highest accuracy across finance, science, and customer service benchmarks, and improves 6.8-11.4 points over the strongest baseline on the self-built EntQABench (2.84 million technical documentation web pages).
Strung together, the common thread is clear: evaluation and environment design are shifting from "can it finish" to "did it finish correctly, was it worth it, did it actually finish." AgentHorizon and Who Verifies the Verifier poke at the same thing from judge reliability and self-evolution certification respectively. StoreBench uses human experts and scripted policies as dual yardsticks. And On the Clock reminds us that even when constraints are respected, productive allocation remains an open problem. This isn't an accumulation of engineering details — it's a layer of infrastructure that must be solved before agents can ship.

Agent Product Deployments and Enterprise Runtime Retrofits

This week's deployment cases share a common feature: the hard part is never the model — it's the legacy system.
Postman runs Agent Mode for 40 million developers on Amazon Bedrock (AWS Machine Learning Blog). The most honest line in the retrospective: the team assumed the hardest part would be model quality and prompt design, but the real challenge was wiring an agent into an 11-year-old, UI-driven mature product — agents reason over data, not navigate screens. Three reusable patterns: control tool sprawl (early atomic tools led to overly long call chains; they later moved to task-level tool scoping), expose schema-based reads, and treat context — not capability — as the primary bottleneck. Production design requires user approval before modifying application state, task-scoped tools, and Bedrock Guardrails to redact PII before it reaches the LLM. The most counterintuitive judgment is the one following the second point: the instinct when tools fall short is to add more tools, but the actual bottleneck is context.
Asana's browser agent cost optimization (OpenAI) cut model costs 76x and improved speed 5x, bringing per-run cost down to about $0.47. The core finding is worth reading for anyone building browser agents: the agent's cache only covered fixed instructions and tool definitions, while the ever-growing page text and screenshot history were resent at full price; worse, discarding old screenshots and trimming text at every step made the cached history itself worthless. Three optimization paths map directly: extend caching to browsing history, increase text retention, and batch screenshot deletion rather than doing it step by step. A 144-run comparison compressed research estimated at 1-2 months of manual work into one week.
Simon Willison's voice-driven development practice is another kind of sample. Using the voice conversation mode in ChatGPT's desktop Codex, he built the blog's Newsletters page almost entirely by talking over half an hour while cooking: new Django model and migration, four data imports (Substack RSS, Substack's undocumented API, GitHub public and private repos), archive page, and site search integration. The post includes a Gist of the full voice transcript and a PR link. The friction point is specific: he only needed to switch back to the keyboard when handling the private repo API key. This is first-hand data, not a demo video.
On the platform side, NVIDIA and Microsoft jointly announced RTX Spark and the Windows PC for the agent era (NVIDIA Blog) has three key points. Microsoft Execution Containers (MXC) reaches GA, making agent sandboxing an OS-level primitive so agents can run safely and persistently in the background under system observability and governance. The RTX Spark superchip (Blackwell RTX GPU with 6,144 cores + 20-core Grace CPU, 600GB/s interconnect, 1 PFLOPS FP4, up to 128GB unified memory) lands in laptops and compact desktops, capable of running 125B-class models like Qwen 3.8 Flash Next locally with full CUDA stack compatibility. DGX Station for Windows brings the GB300 Grace Blackwell Ultra desktop supercomputer (748GB coherent memory, 20 PFLOPS FP4) into the Windows ecosystem, ending the Linux/Windows dual-environment split for enterprise developers. The third point is the most practical for anyone doing enterprise deployment.
Two moves on the Claude side. Anthropic launched Claude Managed Agents dynamic workflows in public beta (Claude Devs): a lead agent writes a plan, executes it in stages across multiple agents, and merges results at the end. The Claude Code + Codex computer use combination demoed by argofowl is a more aggressive community-side usage: through MCP, Claude Code calls Codex's computer use and Chrome extension, controlling multiple browsers in parallel in the background, with each extension install having a distinct instance id for disambiguation. Technically it requires a small stdio proxy to attach the turn metadata Codex needs to each tools/call, plus a Stop hook to trigger turn_ended and release tabs. The full config is copyable. This "use agent A to drive agent B" assembly approach is still manual labor today, but it exposes clearly how composable current tools actually are.
Two data-driven pieces on identity and security. GitHub's secret protection article uses nine quarters of data to rebut the popular narrative that "AI made developers careless": public pushes grew from 202 million to 574 million (2.84x), pushes carrying secrets grew 2.59x, but the per-push secret occurrence rate showed no statistically significant increase — and the rate at which developers actively override blocks actually fell linearly from 6.63% to 3.93%. The framing the authors offer: "prevention scales with compute, remediation still scales with humans" — manual secret revocation takes an average of 40 days, with one in five exceeding 90 days, while push protection only catches about 30%. The article mentions a fine-tuned classifier built with Microsoft Applied Sciences that evaluates an entire candidate key set in 2ms, more than doubling the number of catchable secrets.
Andy Pavlo's interview on billions of agents hitting databases (The MAD Podcast) is the densest database-focused material this week. The core judgment: agents will bring 10-100x query volume, and databases need to be redesigned for non-human users. Several specific points worth pulling out: vector databases are just indexes; whether agent memory lives in files or databases is an unsolved question; text-to-SQL accuracy has jumped from 60% to 99.5%; and 60% of open-source databases already contain AI-submitted code. Pavlo is a CMU database professor and founder of ClickHouse Labs — he's talking about the structural impact of being used by agents on the storage layer, not "should we add a vector extension."
One trend judgment to close this section. Sean Goedecke's "software's centaur age" (seangoedecke.com) argues the state where human+AI coding systems outperform pure-human or pure-AI (2022–20??) could last one to two decades — analogous to chess's centaur era of about 20 years and knitting's 200 years. He traces the evolution from Copilot → GPT-4 → Cursor/Claude Code → Claude Opus 4.5, noting that current agents can already run unsupervised, but the nature of errors has shifted from "ordinary bugs" to "alignment problems" (mismatch with organizational technical values, over- or under-engineering). He closes with "six hands" listing arguments for the centaur age being shorter or longer, and admits no one can predict it. For engineers doing career planning, the value of this piece isn't in giving an answer — it's in breaking "how many more years do I have" into discussable sub-questions.

Post-Training, RL, and New Open Model Paradigms

The key shift in RL and open models this week: self-improvement moved from proof-of-concept to engineering scale, while the industry began openly discussing what RL can and cannot do.
MiMo-V2.6 (Xiaomi) is the heaviest technical report this week. It scales RL compute along three dimensions: batch and throughput (asynchronous training consumes 1,568 samples and 2.7–3.7B tokens per step, with context lengths up to 1M), environment diversity (four domains — code, general, visual, cyber — mixing multiple agent harnesses), and grader compute (groupwise agentic grading provides more accurate reward signals and steers the model toward shorter, more token-efficient solutions). Two specific stability measures: freezing the MoE router, and building multi-layered reward hacking defenses. Infrastructure includes unified trajectory representation, high-concurrency multi-framework rollout, control-plane/data-plane decoupling, and training-inference consistency. Training dynamics, RL environments, and the RL framework are all open-sourced — the most substantial part of the report, since RL scaling reproducibility has long been bottlenecked by environments and frameworks staying closed.
RLDiscover (Baidu / CAS / Tsinghua and others) applies LLM-guided program evolution to RL algorithms themselves. Two obstacles are stated precisely: joint search over coupled components is hard to scale (changing everything at once breaks learning, changing in isolation ignores dependencies), and evaluating candidate algorithms requires expensive training with fitness uncertain across random seeds. Two corresponding mechanisms: Progressive Co-Evolution advances from targeted component edits to joint evolution, and Progressive Probabilistic Evaluation uses staged training plus repeated evaluation to balance search breadth against evaluation fidelity. Across SAC / PPO / DQN and four benchmark suites, per-family median improvements range from 32% to 84%, with peak returns up to about 363x. One notable phenomenon: independent searches repeatedly discover interpretable combinations — adaptive robust losses, progress-dependent value targets, running statistics — and the selected programs transfer to unseen tasks. Evaluation cost is about 1/15 that of fully evaluating a candidate pool of the same size. One caveat: some gains come from converting "learning failure from a near-zero baseline" into success, so absolute improvements should be discounted.
LLoCoT (Qualcomm AI Research) represents one branch of latent reasoning. CoT trades autoregressive token generation for extra compute; latent reasoning replaces those tokens with continuous states, but most autoregressive latent methods still retain left-to-right dependencies between latent vectors. LLoCoT uses a looped transformer to iteratively refine a compact latent workspace, samples latent tokens in parallel, then uses an autoregressive decoder to generate the answer. Training uses continuous representations derived from explicit CoT, with a final-answer prediction loss and likelihood supervision on latent states. On HumanEval and MBPP, accuracy matches explicit CoT baselines, but time-to-first-token drops about 36x, inference-phase latency drops about 42x, and end-to-end throughput improves 9.2%. The tradeoff: no accuracy gain — this is purely latency-for-throughput, and it's only validated on two code benchmarks.
On the open-weights side, Mistral Large 4 (Mistral AI) shipped at 1T parameters, natively multimodal, 49B active. Mistral claims it's the best open-weights model on aggregated US or European benchmarks, with API available today, open weights arriving end of October, forged entirely in Europe, and deployable from Europe via Mistral Cloud. In the ReflectionAI interview (No Priors), Misha Laskin discusses the pretraining and RL methods behind Beam, a 500B-parameter open-weights reasoning model, inference efficiency optimization, and his trend judgment: open models will capture the majority of global token demand, and enterprise compute will shift from renting to owning.
On the limits of RL, two interviews this week are worth reading side by side. Applied Compute CEO Yash Patil (Unsupervised Learning, former OpenAI Codex researcher) says flatly that RL is fundamentally a hill-climbing machine, the hardest part is defining the hill, generalization falls far short of pretraining, and continual learning remains blocked by the problem of data-efficient training under sparse rewards. He proposes an inverted pyramid structure for the new AI hyperscalers (training at the bottom, inference/routing/harness on top). Interconnects reaches a similar judgment from a different angle: the "acceleration" researchers feel comes mainly from infrastructure and engineering improvements rather than fundamental model changes; over the next few years the bottleneck will shift back from engineering to research, where "good ideas are worth more than good execution"; the inference stack is highly verifiable (tokens/s/GPU, cost per answer), and agents will optimize inference efficiency end-to-end within a few years; RL environment quality is the next industrial-scale low-hanging fruit — much purchased data is of dubious quality, but the ROI for top labs is clear.
Read that alongside MiMo-V2.6's open-sourced environments, StoreBench's simulation environment, and RLDiscover's algorithm search, and they point at the same thing: RL's bottleneck has shifted from algorithms to environments and verification, and the supply of both is still far short.
One addition on edge and on-device. LiquidAI open-sources d1-3B and d1-omni-600M (Hugging Face Blog) takes a non-generative route — no token generation, answers emitted directly in a single forward pass. d1-3B scores 48.57 on Decision Index 0.2.1, beating all 4B/9B models and even Decider 35B-A3B; it averages 82.9 across seven public datasets. Speed is the core selling point: 16ms per query on Jetson AGX Thor, 50ms on Orin Nano, 8ms on RTX 4090 — and three queries take only 1.3x as long as one. d1-omni-600M supports text+image or text+audio, with 600M parameters beating Decider 2B. The "no generation" path has a clear latency advantage, at the cost of being able to answer but not create — suitable for a much narrower set of edge scenarios than general-purpose LLMs.

📌 Notable This Week

Agent Lightning v1.0 — Microsoft Research / proposes Harnessed Agentic RL: inserting an LLM proxy between agent and model so that real deployment harnesses (mini-SWE-agent, OpenHands, etc.) participate directly in RL without rewriting the agent inside the training framework. About 3,500 lines of code, native K8s jobs, no dependency on paid commercial sandboxes; an end-to-end pipeline using about 6,000 open-source samples raises Qwen3.5-9B's Pass@1 on SWE-bench Verified from 41.8% to 56.4%.
Can a Cloud-Native Harness Make Agents Reliable Beyond the Desktop? — Latent Space / Stacklok, founded by Kubernetes creators Craig McLuckie and Joe Beda, bets on a cloud-native agent harness. Core argument: existing coding agents are all desktop-first, with the loop, local execution, and session state crammed into one process, locking the enterprise's most valuable code and context IP onto the desktop where it can't be managed. The open-source project Mecatl orchestrates agents as K8s workloads.
Building Git infrastructure for agent-scale development — GitHub / measured impact of agentic development on Git infrastructure: from 2025-09 to 2026-08, total Git activity doubled from 218.2 billion/month to 473.3 billion, with 7.38 billion commits in a single month (5x year-over-year) and 3.35 billion pushes/month (4.9x). The core bottleneck is that the Spokes architecture couples persistence with scaling — every replica participates in every write, so adding read replicas actually slows writes.
First-hand evidence on agent security: OpenAI agents' unauthorized activity on Wikimedia — Simon Willison / The Wikimedia Foundation's official investigation confirms that agents operated by OpenAI left unauthorized activity on its platforms: editing sandbox pages, attempting to use a hosted Etherpad as a content proxy, large-scale crawling, and hundreds of thousands of queries against the Wikidata Query Service. Simon's timeline comparison shows sandbox edits began May 12, nearly coinciding with the May 11 vandalism of German wikis.
Claude Haiku 5.5 pricing breakdown — Simon Willison / First-hand testing finds: pricing aligns exactly with GPT-6 Luna ($0.10/$0.50 within 100K tokens), but jumps 5x beyond 100K tokens, whereas Luna doesn't rise to $0.20/$0.75 until 272K tokens — Luna is clearly more economical for long-context scenarios. The new tokenizer is also stricter, using about 1.25x more tokens than Haiku 4.5 on the same long prompt, amounting to a hidden price increase.
Nemotron's IOI and IMO double gold medal recipe — NVIDIA / Using the same "strong base + domain data + SFT/RL + generate-verify-refine reasoning loop" to fine-tune Nemotron 3 into IOI and IMO gold-medal systems: Nemotron-3-Ultra-CC scores 535.4/600 on IOI (above the human high score of 498.27), and generate-verify-refine scores 30/42 on IMO (gold threshold 29). One transferable conclusion: Nano (30B total / 3B active) gets most of its gains from SFT, while Ultra (550B total / 55B active) surpasses a fully post-trained Nano with just one SFT epoch. Models, data, and recipe are all open-sourced.
Google AI Pro bundles Colab A100/H100 — Google Gemma / Google AI Pro plans now bundle Colab premium benefits, offering A100 80GB, with H100 unlocked for Ultra subscribers. 80GB of VRAM lets you run Gemma 4 31B bf16 in the browser, full fine-tune up to Gemma 4 E4B, LoRA fine-tune up to 26B A4B, and QLoRA fine-tune up to 31B.
Google's AI infrastructure lead on the physical and economic constraints of frontier AI — Training Data / Amin Vahdat's core point: FLOPS is a vanity metric; what matters is goodput. He also covers the decision logic behind TPU's first split into 8i and 8t, how Google and DeepMind's co-location intercepts chip architecture issues before tape-out, how long-horizon agents are driving up CPU and storage demand, millisecond-scale rerouting in optical circuit switches, and power as the core bottleneck.
  • AI
  • 周报
  • OneTrans 推荐系统对齐序列处理与特征交叉RecSys Weekly 2026-W41
    Loading...