- 标签:
- AI (228)
- Daily (201)
- Tech Trends (201)
- 周报 (30)
- Recommendation Systems (25)
- Weekly (25)
- Papers (25)
- 推荐系统 (16)
- 思考 (6)
- 论文 (6)
- Agentic Engineering (6)
- 日报 (5)
- 技术趋势 (5)
- 深度学习 (4)
- Harness Engineering (3)
- 推荐 (2)
- 工具 (2)
- 强化学习 (1)
- 思维模型 (1)
- Transformer (1)
- LLM (1)
- 管理 (1)
- 生成式 (1)
Open-source AI hit a new inflection point: Z.AI revealed the mysterious Ox Alpha as GLM-5.3-Flash — a 320B-A18B MoE with MIT license, trained at 1/9 the cost of Qwen3.7-Plus, and already proven on domestic Chinese AI chips with 3x inference efficiency gains. Alibaba countered with Qwen3.8-Flash (125
AI infrastructure hit a turning point today. NVIDIA unveiled the Vera Rubin NVL72 with 30x efficiency gains over GB300, while Meta countered with its own MTIA 300 training chip and MetaRoCE network protocol — the full-stack war is on. Hugging Face is exploring a $13B sale, nearly tripling its 2023 v
The "free lunch" era for AI agents is officially over. Anthropic's flagship Fable 5 model is struggling with adoption — just 8% share per Ramp's index — while the company's annualized revenue hits $65B. Drew Breunig's analysis crystallizes the shift: teams now route work by model tier, using GLM 5.2
AI hit a major milestone today: NVIDIA's AVO architecture scored a perfect 100% on ARC-AGI-3, proving system design — not just model scale — can unlock frontier-level performance. Anthropic is reportedly prepping a $100B+ IPO that could top SpaceX's record, with a $2T valuation target. Meta struck b
This week in AI, one clear theme dominates: the accelerating pace of capability gains is forcing evaluation, governance, and safety systems to speed up in tandem. Sam Altman announced a pause on some frontier RL training — the week's biggest industry event — citing capability advances outpacing the cadence of safety and alignment work. Meanwhile, NVIDIA's AVO agent completed all 183 tasks on ARC-AGI-3 with a 100% score, and Ornith-1.5 matched Claude Opus 4.8 through self-improvement training. Capability and safety are both accelerating — and pulling against each other. The second thread is a paradigm shift in evaluation and training environments: from static, zero-shot benchmarks toward real-time, long-horizon, self-generated environments. This week's FM-Bench (AnalogyAI) tests long-term decision-making via 20 years of football management simulation; EnvHarness (Google) makes static environments adaptively evolve; Wuying-Browser-Agent (Alibaba Cloud) introduces BrowserBench with an average of 37.9 steps. Evaluation is no longer asking "can it do it" — but "can it do it reliably over the long run." The third thread: inference optimization has entered a phase of fine-grained engineering. LFM2.5-DSpark (Liquid AI) delivers 3.2x speedup via speculative decoding; LMSYS's weight-caching daemon cuts engine loading from 495 seconds to 0.63 seconds. The cost curve is being pushed down on multiple fronts.
AI hit a major inflection point today: Z.ai CEO Tang Jie declared "parameter count is dead," crediting GLM 5.3's leap entirely to long-horizon RL training in synthetic environments. NVIDIA dropped $6B on Poolside's "model factory," while Moderna/Merck's personalized mRNA cancer vaccine hit Phase III
AI safety took center stage as OpenAI paused frontier RL training, with Chief Scientist Jakub Pachocki signing the Pacing the Frontier initiative. China's Z.ai shipped GLM-5.3 — matching Kimi K3 on intelligence while costing 19% less per task — and DeepSeek faced scrutiny after independent tests sco
AI infrastructure hit a consolidation milestone: Stripe acquired OpenRouter for $7B, just 90 days after its $1.3B Series B — a clear signal that the model routing layer is now a strategic chokepoint. Meanwhile, NVIDIA committed 4.25 gigawatts of AI factory compute to OpenAI under a 20-year, ~$600B L
AI hit a commercialization inflection point: Anthropic posted its first adjusted operating profit with $11.5B in Q2 revenue, while OpenAI's enterprise income overtook consumer for the first time. The open-source world pushed back hard — DeepSeek raised API prices, prompting OpenCode's "Operation Che