type
Post
status
Published
date
Sep 26, 2026 05:41
slug
ai-weekly-2026-W39-en
summary
One keyword this week: cost per task. On September 22, Anthropic released Claude Opus 5.5, running 40% cheaper than Opus 5. About an hour later, OpenAI released GPT-6 Sol and Luna, with API prices cut in half from GPT-5.6's promotional pricing. The same week, StepFun shipped Step 5 Preview (600B/27B MoE, $0.71 per task), and Xiaomi trained MiMo-V2.6-Pro — a 1T total / 42B active open-weights model — for roughly $3M. These launches are no longer about "who's smarter." They put intelligence and cost on the same Pareto chart. In Artificial Analysis's evaluation, GPT-6 Sol's cost per task dropped about 50% versus the prior generation, while generating *more* tokens per task — the savings come entirely from unit price. The second thread runs on the inference side, where two directions compress the bottleneck at once. One is System 1 decision models: Stanford's CLM-8B uses contrastive learning to connect states to actions, running 9× faster than Jev at 81.6% on DeepSWE; LMSYS built multi-candidate scoring for Jev-class models on SGLang, cutting 16-candidate p95 from 54.1ms to 20.6ms. The other is KV cache quantization: NVFP4 on Blackwell compresses per-token KV to 56% of FP8, speeds up 1M-context decoding by 78%, and stays near-lossless on GPQA and AIME. OpenAI's GPT-6 prompt caching update, shipped the same day, attacks the same problem from the more application-layer angle of cache hit rate. The third thread is agents moving from "it runs" to "it's managed." Accenture offers an enterprise harness routing scheme that recovers 14–21% of model spend in a 10,000-seat simulation. Microsoft's LIMBO sandbox uses 25,930 episodes to pull apart where exactly-once semantics should live — the model, the harness, or the tool contract. Nubank screens models via simulation on a product serving 140M customers, lifting online tNPS by 36.69 points.
tags
AI
周报
category
AI Tech Report
icon
password
priority
1
📊 Weekly Overview
One keyword this week: cost per task.
On September 22, Anthropic released Claude Opus 5.5, running 40% cheaper than Opus 5. About an hour later, OpenAI released GPT-6 Sol and Luna, with API prices cut in half from GPT-5.6's promotional pricing. The same week, StepFun shipped Step 5 Preview (600B/27B MoE, $0.71 per task), and Xiaomi trained MiMo-V2.6-Pro — a 1T total / 42B active open-weights model — for roughly $3M. These launches are no longer about "who's smarter." They put intelligence and cost on the same Pareto chart. In Artificial Analysis's evaluation, GPT-6 Sol's cost per task dropped about 50% versus the prior generation, while generating *more* tokens per task — the savings come entirely from unit price.
The second thread runs on the inference side, where two directions compress the bottleneck at once. One is System 1 decision models: Stanford's CLM-8B uses contrastive learning to connect states to actions, running 9× faster than Jev at 81.6% on DeepSWE; LMSYS built multi-candidate scoring for Jev-class models on SGLang, cutting 16-candidate p95 from 54.1ms to 20.6ms. The other is KV cache quantization: NVFP4 on Blackwell compresses per-token KV to 56% of FP8, speeds up 1M-context decoding by 78%, and stays near-lossless on GPQA and AIME. OpenAI's GPT-6 prompt caching update, shipped the same day, attacks the same problem from the more application-layer angle of cache hit rate.
The third thread is agents moving from "it runs" to "it's managed." Accenture offers an enterprise harness routing scheme that recovers 14–21% of model spend in a 10,000-seat simulation. Microsoft's LIMBO sandbox uses 25,930 episodes to pull apart where exactly-once semantics should live — the model, the harness, or the tool contract. Nubank screens models via simulation on a product serving 140M customers, lifting online tNPS by 36.69 points.
Frontier Models: Halved Costs Become the Main Battlefield
September 22 was dense. Anthropic went first with Claude Opus 5.5, the first model in the new Claude 5.5 family. The company says it reaches Fable 5.1 levels on most tasks while running 40% cheaper than Opus 5. About an hour later, OpenAI released GPT-6 Sol and Luna.
GPT-6 Sol/Luna is the cost-efficiency tier after Astra: same training methods, capabilities carried over in professional work, factuality, coding, computer use, and alignment — but API prices cut in half. Sol goes from $4→$2 input and $20→$10 output; Luna from $0.20→$0.10 and $1.20→$0.50. Sam Altman's tweet frames it as "half the price per token, lower price per task." OpenAI attributes the cuts to caching and inference-side improvements — they published a piece on better prompt caching the same day, with mechanism details in the next section. Both models are confirmed via the official ChatGPT account as rolling out in ChatGPT Work and Codex for Plus/Pro/Business/Enterprise/Edu.
Artificial Analysis's itemized evaluation gives a clearer picture. On cost: GPT-6 Sol (max) spends $1.06 per task on the Intelligence Index, about 50% cheaper than the prior generation's $1.99; Luna (max) costs $0.07 per task, about 60% cheaper. But both models actually generate *more* tokens per task (Sol 31k vs 29k, Luna 51k vs 41k) — the savings come from unit price, not efficiency. On coding: Sol scores 57 on the Coding Agent Index under the Codex harness (+2), Terminal-Bench 4.0 goes from 37% to 43%, SWE-Atlas-QnA from 54% to 58%, with cost per task at $2.99, roughly halved. Luna actually regresses 2 points (SWE-Atlas-QnA 44% vs 49%). Hallucination narrows across both generations, but the mechanism is worth noting: Sol's hallucination rate drops from 92% to 60%, at the cost of answering only 83% of questions (down from 99%), with accuracy falling from 59% to 54%. On GDPval-AA v2.1, Sol loses about 100 Elo — the clearest regression this round. After manually reviewing outputs, the evaluators attribute the drop mainly to degraded presentation quality and deliverables missing rubric elements.
Simon Willison's writeup adds two details: GPT-5.6 will go back up 25% in November, so the real reduction against the official comparison baseline is larger; and Opus 5.5, at max effort, failed for the first time to return the "pelican riding a bicycle" SVG. The latter is trivia. The former shows the comparison points on the pricing page were chosen with care.
Chinese labs played their hands the same week. StepFun first introduced Step 5 Preview in a tweet: 600B total / 27B active MoE, 1M context plus vision, aimed at agentic and professional work with particular emphasis on financial scenarios. Official numbers: AA index 44, $0.71 per task. Open weights on October 15. Read alongside the KITE architecture paper in the next section, part of this model's capability-cost curve adjustment happened at the architecture layer.
Xiaomi is more noteworthy. MiMo-V2.6-Pro is a 1T total / 42B active open-weights model, with training cost reported at roughly $3M, shipped alongside Flash and UltraSpeed (up to 20× output speed at equal quality). The technical report lays out RL along three axes: fully asynchronous large batches (1,568 samples per update, up to 1M context, 3.5–3.7B tokens per step), multi-task multi-harness environment mixing, and group-relative grader compute providing long-horizon rewards. Xiaomi also open-sourced the environment code and training recipes — five sets covering code, ARVO vulnerability reproduction, general, webdev, and music, plus composable mini-harness configs. The 7k+ task dataset is not yet released. Team lead LuoFuli said in a tweet this is "one of the largest single RL compute runs by an open-model team to date," and additionally released a Qwen model distilled from MiMo RL trajectories, 7K environments, and the full RL framework — the rationale being "let the community focus on real Agentic RL problems instead of building environments from scratch." He also offered a training-strategy judgment: use MixRL on verifiable medium-difficulty tasks (code and agentic), and train hard-to-verify, ultra-long-horizon, or purely subjective-signal tasks separately with MOPD before merging.
Taken together, this week's cost reduction happened at multiple magnitudes: frontier labs compress unit price with caching and inference infrastructure, second-tier labs compress cost per task with architecture and MoE activation ratios, and the open-source side trades RL compute for training cost. Teams doing model selection need to redo the math.
System 1 Decision Models and KV Cache: Two Faces of the Same Bottleneck
Both inference-side threads this week point at the same observation: in agent scenarios, the LLM bottleneck is often not "thinking too shallowly" but "deciding too slowly, caching too expensively."
Decision models first. A Stanford team (Azalia Mirhoseini's first-hand post and first author jackyk02's thread) introduced the Contrastive Language Model (CLM), which connects states and actions directly through contrastive learning, skipping the autoregressive token-generation step. CLM-8B matches Jev's zero-shot performance on computer use, gaming, and tool calling while running 9× faster; after lightweight fine-tuning it hits 81.6% on DeepSWE and 87.6% on Terminal-Bench 2.1. The team also established scaling laws across compute, model size, and data size, and open-sourced checkpoints, data, and training/serving infrastructure. For agent practitioners, this is a viable starting point for pulling "decision" out of the generation task.
The design philosophy behind this thread is laid out most clearly in Latent Space's Jev interview. Diogo Almeida is an InstructGPT co-author and now CEO of TypeSafe AI. His core claim: after ChatGPT's success, the industry was left with one path — autoregressive chat fine-tuning — which dragged alignment, refusal, and reliability off course; none of the three RLHF variants is the right north star. He argues for small, production-oriented System 1 models, stresses that KV cache determines everything, and says models should fade into the background like regex. The interview covers Jev's deployments in coding agents, voice plus browser control, entity resolution, and natural-language search, and previews the ReasoningJev direction.
On the deployment side, LMSYS offers the engineering approach. They serve Jev-class decision models with SGLang: the problem is that generate plus top-k logprobs may drop the needed label, and shared context gets recomputed for every candidate. SGLang's two countermeasures are `/v1/score`, which returns scores for specified labels directly, and Multi-item scoring (MIS) — shared context computed once, each candidate isolated. The result: MIS latency stays nearly flat from 2 to 16 candidates, and 16-candidate p95 on Qwen3-8B drops from 54.1ms (Generate) to 20.6ms (MIS); on Qwen3-0.6B under rising load, MIS p95 stays under 100ms while Generate and SIS are already in the seconds.
The other thread is KV cache quantization. LMSYS, working with Alibaba Qwen and NVIDIA, implemented NVFP4 KV cache on Blackwell: per-token footprint is about 56% of FP8, decode throughput improves 37% / 58% / 78% at 32K / 160K / 1M context respectively, GPQA-Diamond and AIME 2025 match FP8 on Qwen3.5-397B-A17B, and capacity at 1M context is 1.78× FP8. The implementation combines three components: NVFP4 two-level scaling, in-kernel dequantization during decode, and paged KV cache; enabling it in SGLang takes one flag, `--kv-cache-dtype nvfp4`. Where AgentX throughput turns over under FP8, higher cache hit rates let NVFP4 keep scaling linearly.
At the architecture layer, StepFun's KITE paper offers a more aggressive idea. It targets simultaneous reduction of training cost, prompt processing, and decoding by placing new parameters — when upcycling from small to large models — in regions that don't affect attention KV. The concrete implementation is SST (Step Scale Transformer), a dual-tower decoder where one tower produces KV and the other reads it. At equal cumulative training compute, SST is a 67B MoE activating 2.15B body parameters per decode token, with training loss below 47B/63B MoE controls at 1.48B/2.02B active, and estimated inference cost reductions of 6.7% and 31.6% respectively. The experiments only compare within a single model family, include no downstream benchmarks, and release no code — so read it as a signal about architecture direction, not a reusable solution.
Training-side efficiency has an independent line too. Amazon's CounterRoute paper tackles "when should the model think versus answer directly." Its approach jointly learns routing and mode-conditioned responses on a shared policy, uses counterfactual rollouts from the current policy to assign cross-mode credit only to routing tokens, trains response tokens with within-mode GRPO, and stabilizes early training with a paired-to-self-routed curriculum. Across nine benchmarks, macro-average accuracy improves while Qwen3-8B cuts generated tokens by 51% and Qwen3-14B by 41%; on instruction-following and commonsense benchmarks where direct answering is stronger, think rate can drop to 1% with better quality. Training covers only math and instruction following, but routing behavior and response quality generalize to coding, science, knowledge, and commonsense.
Finally, back to the practical layer: OpenAI's GPT-6 prompt caching update is this week's most directly impactful piece on long-horizon agent costs. Reusing a shared prefix within a 30-minute window earns up to 90% off input tokens; a new Prompt Caching Dashboard tracks hit rate and input composition; cache misses return structured JSON (e.g. `reason=tools_changed`, affected token count), pinpointing whether the model, tools, settings, or input broke reuse. Three engineering practices to preserve caching: use explicit cache breakpoints to choose the cached prefix; on GPT-6 you can adjust reasoning effort mid-conversation without breaking cache (append a `configuration_update`); keep tool definitions/schemas/order stable, use `allowed_tools` or `tool_choice=none` instead of deleting definitions, and append new instructions to the end of context. GitHub Copilot measured a drop of over 50% in missed-token share.
Put these four together: System 1 decision models change *whether to generate*, KV cache quantization and KITE change *how expensive generation is*, and CounterRoute plus prompt caching change *when it's worth paying*. Agent inference routing, caching, and quantization are being split into independently optimizable layers.
Agent Harness Governance and Enterprise-Scale Deployment
This week's most interesting technical papers all circle one very concrete question: when enterprise agents ship, where does the money go, where do the errors come from, and who's responsible.
Accenture Responsible AI's Control the Harness, Control the Cost starts from cost structure. The paper's premise: most enterprises don't build their own harness — they buy Anthropic Claude Code or OpenAI Codex. But the harness decides which model answers, what the model reads, how prompt cache is used, and which subagents run. It picks the rate on the price sheet *and* determines how much volume you buy at that rate. Enterprises that keep harness defaults inherit those choices and the bill. The authors build a customizable router that labels each prompt against the enterprise's own taxonomy using Jev (a classifier with calibrated probabilities). Since one user turn is actually multiple requests against the same prompt cache, the router only switches models when no conversation needs cache rebuilding — session start, side lanes, subagent launch. One counterintuitive conclusion derived from the price sheet: in long tool-heavy sessions, the most expensive model can be cheaper than the next tier down (reproduced across roughly 10,000 real sessions). In a 10,000-seat simulation, it recovers 14–21% of model spend ($3.3M–$5.0M/year) at Anthropic's public prices as of September 21, 2026. The paper also surveys risks across 20 harnesses, single-vendor dependency pricing, and offers an enterprise-runnable control plane plus a decision ladder for "should we build our own harness."
Microsoft's LIMBO paper handles the reliability side: when a tool call times out or returns a server error, the operation may already have taken effect — blind retries mean duplicate charges, duplicate posts, duplicate deploys. The core question is which layer should enforce exactly-once semantics: the model, the harness, or the tool contract. The authors built a deterministic sandbox, LIMBO (6 services, 12 failure modes, ledger-based scoring), and ran 25,930 episodes across 9 recent models, 3 production harnesses, 2 contract variants, and 15 recovery conditions. Conclusions split by failure type: when read-back is immediate, the model decides — frontier models instructed on exactly-once almost never duplicate writes (0.5%), weak models duplicate often, and the model explains 53% of explained variance. When read-back isn't possible (request still in flight, or transport-level duplicate delivery), the same frontier models duplicate at 56% and 74%, and the contract explains 81%. The paper also proves that without an in-flight time bound, no verification-only strategy can guarantee exactly-once under late commit. Waiting only works when in-flight latency has a short, known bound; under heavy-tailed latency, even waiting an hour per episode loses to giving each write an idempotency key — duplication drops from 28% to 4%, because agents use the key when they see it. The harness barely affects outcomes, and a keyed guard transfers seamlessly across harnesses. One easily missed signal: agents self-report success in 90% of duplicate-effect scenarios.
Nubank's Screen Before You Serve pulls the view up to production at 140M-customer scale. Improving CX agents is hard: they must recognize intent, follow complex operational policy, and use tools reliably, while manual end-to-end testing has limited coverage and online experiments expose customers to failure. The authors use a hypothesis-driven simulation workflow (the Snowglobe simulator) to screen candidate agents before deployment: synthetic customers react to agent responses, and simulated tool outputs let multi-step agentic workflows run without hitting production backends. On the Card Delivery agent and its successor Card Management (Brazil's highest-traffic chat-support agent), simulation correlates strongly with production on version-level binary evaluation scores across 4 deployed versions; simulation-guided iteration lifted tNPS by 36.69 points (online A/B). They also used 16,000+ simulated conversations to screen open-weights model configurations, and a subsequent online A/B raised self-service rate by 8.82 percentage points to Nubank's all-time high, with no statistically significant change in tNPS.
On the training side, AWS AI / Amazon's PSP paper diagnoses a common but hard-to-spot problem: on-policy self-distillation (OPSD) teaches students in multi-turn agents to act *confidently* but doesn't teach them where information comes from; the trained agent behaves as if it holds privileged information it never observed, with worst-case performance below the untrained base. The authors propose Privileged Self-Practice: move privileged information out of the loss and into the sampler. Concretely, when rollouts mostly fail on a task, inject a per-task instruction written by an analyzer model, resample with the instruction, and train with an unmodified GRPO objective. On AppWorld and SWE-bench Verified, three different student models all achieve their best average scores, and it's the only method that consistently beats plain GRPO — up to +65% task goal completion on AppWorld and up to +61% resolved rate on SWE-bench Verified.
Alibaba Tongyi's Qwen-Planner-Agent is a typical case of systematically integrating these ideas. The paper builds a closed-loop AI-for-AI framework: AI for Data uses specialized agents to create tasks, sample trajectories, and filter balanced data, then feeds training feedback back into subsequent data generation; AI for Training uses supervised planning cold-start plus mixed-environment online agentic RL, where CARE (Competence-Aware Reward-and-Advantage Engineering) cuts inference and tool-call costs while preserving task performance; AI drives model-harness co-evolution, orchestrating memory, skills, and tools at execution time and feeding structured action feedback plus retained failure trajectories back into joint model-harness adaptation. Results: best among all evaluated models and systems on the in-house MobilePA-Bench (77.1), beating the base across all four dimensions — tool use, memory, skills, and subagent coordination. The limitations are clear too: it relies mainly on an in-house benchmark, and non-mobile benchmarks are reported only as "improved" with no specific numbers.
Read these five together and the enterprise-agent conversation has shifted levels — from "which model is best" to "where does routing live, who installs the idempotency key, how do we screen via simulation, does privileged information go in the loss or the sampler." None of these are solvable at the model layer. They require joint design of harness, tool contract, simulation, and data.
Coding Agent Sandbox Isolation and End-to-End Automation
Two pieces this week on agent runtime environments: one on why isolation is mandatory, one on what becomes possible once responsibilities are separated.
Package Manager Sandboxing is the week's most counterintuitive read. It starts with Homebrew 7.0.0: installation splits into a networked fetch phase and an offline install phase, replaces Bubblewrap with kernel-level Landlock, and migrates arbitrary Ruby `post_install` hooks to signed declarative `*_steps`. The author then points out that this sandboxing primitive stack (Seatbelt / Landlock / seccomp) is exactly what Claude Code, Codex CLI, Cursor, and Gemini CLI use — for a straightforward reason: both are running unreviewed code. More importantly, the two have already crossed paths in attack and defense: agents install packages for you; the Shai-Hulud worm persists by writing into Claude Code config via postinstall scripts; Mandiant reports attackers hijacking coding-assistant sessions to spread across roughly 100 internal repos. The article offers a gating / confining / replacing taxonomy and names the unsolved problem both paths share: how to verify the output of a confined phase. It doesn't give answers, but it makes the shared lineage of two seemingly unrelated fields clear.
GitHub Security Lab's Fuzzing Taskflow Agent is the opposite kind of sample: hand the entire fuzzing pipeline for C/C++ projects to an LLM agent — automatically identify entry points, analyze the build system, write harnesses, run AFL++, read coverage, iterate on harnesses, triage crashes, and generate vulnerability reports. The architecture draws a clear separation of duties: the agent only makes decisions, MCP tools expose only primitives like `run_afl_for` / `compile_harness`, and all state lives in SQLite. Each harness builds two binaries (`.afl` for fuzzing, `.cov` for real line/branch coverage replay), forming a coverage-feedback loop. The article explicitly warns: this pipeline executes LLM-chosen build commands directly on the host, so it must run in a disposable environment. This "agent decides + MCP primitives + separate state store" layering is one of the week's reusable agentic pipeline design patterns.
When chat is the wrong UI enters from the interaction paradigm. The core argument: chat is the LLM's default interface, but for most task scenarios it's the wrong one. The author uses the GitHub Copilot app's canvas as an example — a full-stack mini-app running inside the app with no browser shell, capable of bidirectional communication with the agent, calling third-party APIs, and executing code locally. The key insight: rather than having the agent repeatedly perform operations like `stage and commit` and burn tokens, have the agent build a tool so all subsequent interactions are free. The piece walks through Connect 4, Winget package management, a SQLite browser, and a Jekyll editor, showing how canvas takes you out of the loop in the research→prototype→plan→implement→iterate→finalize workflow.
One last product note from Microsoft and OpenClaw. Microsoft released Autopilot, a persistent agent built on OpenClaw, led by Omar Shahine's team, with OpenClaw officially emphasizing that Microsoft "contributed a large number of improvements back to OpenClaw." The news itself carries limited information, but placed alongside the sandboxing and canvas items above, it shows coding agent product form growing toward "persistent, multi-entry, third-party buildable."
Biosecurity and the Industry Safety Debate
This week's "safety" topic has two opposing narrative threads.
One comes from Latent Space's interview with Radical Numerics CEO Eric Nguyen. He tells AI biosecurity as a complete technical arc: he long pushed Genomic Language Models at Stanford without acceptance from the biology community, eventually contributing to Evo / Evo 2 — models that the Arc/Stanford team has used to generate whole bacteriophage genomes and synthesize functional viruses. The interview explains what makes DNA language special (4-letter alphabet, 60K for a single gene, 3B for the human genome) and why long-context innovation (StripedHyena) is a prerequisite for biological intelligence. One experiment stands out: given only low-scoring aptamers, the model extrapolates along the score trajectory on its own and reproduces unseen high-scoring sequences — effectively "chain-of-thought in DNA." His core position: models that boost biological capability also help defense keep up, and defense is currently losing, so push the frontier harder. Read alongside the earlier OpenAI→HuggingFace attack incident, this is a rare public discussion of biosecurity — one of the two risk domains frontier labs explicitly name — with actual technical detail.
The second is NVIDIA CEO Jensen Huang on the Ezra Klein Show. He responds to AI safety controversies, regulatory necessity, and technological optimism from an industry perspective, touching on compute, model evolution, and the global AI competition landscape. His core argument is that AI doomerism is overblown. The interview's value is understanding how a top industry leader thinks strategically — especially given NVIDIA's position in compute supply, its stance on regulation is itself an industry signal.
The third is Zvi's section-by-section close read of the Claude Opus 5.5 system card. He skips what repeats earlier cards and focuses on the genuinely informative sections: classifier fallback strategy, RSP evaluation (still CB-1, not CB-2), a major policy change in biological evaluations (no longer testing helpful-only versions, switching to a test design that avoids refusals), agentic safety, prompt injection risk, alignment and honesty evaluations, white-box analysis, sandbagging, CoT controllability, and more. For practitioners who want to know which parts of a system card are real signal, this is a high-density guide. One detail worth noting: Anthropic dropping helpful-only versions in biological evaluations reflects test design shifting toward "avoid refusals contaminating evaluation results."
Read all three together and it's one debate from three angles: the technical arms-race argument, the industry-side optimistic response, and the frontier lab's own published safety evaluation framework. The debate has no convergence point, but this week's inputs are more concrete than usual — a synthesizable bacteriophage genome, a discussable industry position, and system card clauses you can compare against.
📌 Notable This Week
OpenRouter: from Seed to Stripe — Latent Space / OpenRouter CEO Alex Atallah and AMP's Anjney Midha recount the journey from betting in the Llama/Alpaca era that "no single model will win" to becoming a neutral routing layer with 10M developers processing over 10 trillion tokens daily, and the full acquisition by Stripe. Worth noting: the claim that "agentic fraud will become the AI economy's defining security problem" — autonomous agents attacking token flows.
Who Feeds the GPUs? Inside AI's Hidden $30B Layer — The MAD Podcast / VAST Data founder Renen Hallak on the overlooked data layer in the AI stack: the DASE shared architecture, why traditional databases collapse under trillions of vectors, infrastructure differences between training and inference, and KV cache, RAG, agent memory, and agent identity/permission security. The shift in customer demand signals from 500PB to 2EB is the episode's most useful anchor.
I Saw the Signal of Scaling Laws | In Conversation with Tsinghua IIIS Assistant Professor Xu Mengdi — Crossroads / Xu Mengdi on robot in-context learning research, discussing the signals and pitfalls of embodied intelligence scaling laws, pretraining ceilings, the data closed-loop problem, and world model trends. Combined with the Generalist GEN-1.5 release, his judgment is that robotics is currently at roughly the GPT-1 stage.
NarrateAI: production-ready LLM quality assurance on Amazon Bedrock — AWS / five-layer QA engineering for a real-time conversational BI assistant: routing by total retrieved passage volume (about 90% take the single-concatenation fast path), cross-account multi-model failover, real-time streaming evaluation, a composite evaluation framework, and two-stage cascaded numerical hallucination checks. Overall it reaches about 99% numerical accuracy under streaming responses.
Runway's WorldPrompt and the Engineering of Real-Time Worlds — Latent Space / exclusive interview with Runway CTO Kamil Sindi and chief research scientist Robin Kahlow, unpacking WorldPrompt, a new GWM Worlds 2 feature — an input format that can pin the first frame and generate timestamped events. Includes a side-by-side comparison table of Runway, Genie 3, Odyssey-2 Pro, and World Labs RTFM, and breaks down the two big engineering problems in turning a video model into a real-time runtime: whole-clip generation to frame-by-frame, and generation down to real-time frame rates.
Offloaded inference for real-world physical AI robotics — Microsoft Research / the first systematic measurement of onboard vs edge vs cloud inference on mobile manipulation tasks: under a lightweight GPU, mapping and planning slow by up to 383%, timely navigation obstacle detection drops 30%, and VLA accuracy falls 50%. The Physical AI Toolchain simultaneously adds Kubernetes-based containerized deployment and distributed inference orchestration.
Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war — Simon Willison / documents the 9/22 model launch wave with a full price comparison table, and notes GPT-5.6 will rise 25% in November, so the real reduction against the official comparison baseline is larger than the announced numbers. Opus 5.5's cache read price drops 60%, significant for long agentic conversations where 90%+ of input goes through cache.