type
Post
status
Published
date
Oct 5, 2026 05:00
slug
ai-daily-en-2026-10-05
summary
Attention architecture is quietly being rewritten. A deep config.json teardown shows mainstream open models have nearly abandoned full attention — Kimi K3 runs 69 linear layers out of 93, DeepSeek V4 has zero full-attention layers, and GLM-5.3-Flash mixes 34 linear with 11 sparse. Meanwhile Google's
tags
AI
Daily
Tech Trends
category
AI Tech Report
icon
📰
password
priority
1
📊 Today's Overview
Attention architecture is quietly being rewritten. A deep config.json teardown shows mainstream open models have nearly abandoned full attention — Kimi K3 runs 69 linear layers out of 93, DeepSeek V4 has zero full-attention layers, and GLM-5.3-Flash mixes 34 linear with 11 sparse. Meanwhile Google's Argon agent autonomously freed 300TiB of memory across data centers, and Anthropic's red-team eval put GLM-5.3 at 4% full control-flow hijack — closing in on Claude Mythos's 6%. On the money side, Intel doubled this year on CPU demand spillover, while Nvidia's depreciation debate and $3.8T in off-balance-sheet hyperscaler commitments stayed the quiet risks.
🔥 Trend Insights
- Full attention is dying: Config teardowns show Kimi K3, Qwen3.8, DeepSeek V4, and GLM-5.3-Flash all dropping full attention for linear/sparse hybrids — long-context model selection now needs a new checklist.
- Agents run production infra: Google's Argon agent autonomously optimized data center memory and freed 300TiB, a rare quantified sample of agents delivering real fleet-level value, not demos.
- Open models close the attack gap: Anthropic's eval puts GLM-5.3 at 4% full control-flow hijack vs Claude Mythos's 6%, with defenses bypassed in 64–100% of tests.
🐦 X/Twitter Highlights
📈 热点与趋势
- GLM-5.3 在 ExploitBench 解出 12% 漏洞利用任务,Claude Mythos 为 14% - 数据来自 Anthropic;文中称花 20.40 美元 token 就找到 Google Chrome 的一个近期漏洞。Andrew Ng 认为这是工程问题,防守方长期占优 @DeepLearningAI(DeepLearning.AI,吴恩达创办的 AI 教育机构)
- Grok 4.7 在 Frontier v4 各努力档位登顶 - 超过 GPT-6.1 Sol、Opus 5.5 和 GPT-6 Astra @morganlinton(AI 内容博主)
- 特朗普政府设立 "Super Intelligence Force" - 任务是评估 AI 风险,并建议联邦政府应扮演什么角色 @Polymarket(Polymarket,预测市场平台)
- Sam Altman:世界应为 AI 收益接受一些坏事发生 - 他在 POLITICO 新栏目 Decoded 谈取舍、AI 安全,以及 OpenAI 与 Anthropic 的差异 @politico(POLITICO,美国政治新闻媒体)
🔧 工具与产品
- NVIDIA 开源 PixelUMM 的代码与权重 - 这个统一多模态模型把 VAE 和 ViT 整个移除,图像与视频的理解和生成都在原始像素空间完成,一个无编码器模型同时做两类任务 @CongWei1230(PixelUMM 作者之一)
- Empero 发布开源 agent Homebrew,自动帮你炼模型 - 用户只说想要什么,它负责挑基座模型、准备数据、设定安全参数并跑完训练 @EmperoAI(Empero,AI 初创公司)
- Cloudflare Web Search API 进入 beta - 通过 AI Gateway 用实时网页数据给 AI 回复做依据,零数据保留 @CFchangelog(Cloudflare Changelog 账号)
- Perplexity Decisions API 一次通关 Pokémon FireRed 四天王与冠军 - 中位响应 592ms,p95 987ms,96.4% 的响应在 1 秒内,137 次实时调用成本约 0.028 美元 @AravSrinivas(Aravind Srinivas,Perplexity CEO)
⚙️ 技术实践
- BF16 FlashAttention-3 训练中梯度范数暴涨 1000 倍,loss 比 FP32 attention 高 0.2 nats - 训练长时间看起来健康,随后 grad norm 跳升,全程没有出现一个 NaN;作者定位问题出在 attention 反向 @Chen94751623484(该研究作者)
- Claude Code 可调用 Codex 的 computer use 与 Chrome 扩展,后台同时操控多个浏览器 - 走本地 MCP server,不用每次弹 "allow" 授权;每个扩展实例有独立 instance id,可区分 chrome 与 helium。同一套 8 任务测试中,Opus 5.5 跑在 Codex 引擎上得 6/8,Codex 自己 6/8,cua driver 只有 3-4/8 且成本 4 倍 @argofowl(社区开发者)
- Triton GPU 编程系列首讲上线,一小时动画讲 H100 内部结构 - 覆盖 SM 与 warp scheduler、tensor core、SRAM/DRAM 与 HBM 的层级差异、coalescing 与 kernel fusion、roofline,以及 FlashAttention 要解决什么 @Mayank_022(GPU 编程教程作者)
⭐ Featured Content
主流开源模型已几乎放弃全注意力:八种注意力设计的成本横评 | 从各家 config.json 反推 2026 架构真相
作者逐行翻查 2026 年 10 月 1 日的 config.json,发现主流开源模型已几乎不再做全注意力:Kimi K3 93 层中 69 层线性 + 24 层全注意力,Qwen3.8 为 69+23,DeepSeek V4 干脆没有全注意力层(只精确读最后 128 token,更早的靠压缩摘要 + indexer 选择),GLM-5.3-Flash 为 34 线性 + 11 稀疏(预算 2048)。文章按「每层还读多少 token」把八种注意力设计排序,逐一说明各自省下什么、放弃什么、谁在出货、该测什么,并给出 DeepSeek V4 省 90% KV cache、27% 每 token 算力等量化数据。一个实用警告:GLM-5.3-Flash 把稀疏层列在 full_attn_layers 键下,要看 layer_types 邻居字段而非标签。对选长上下文 / agent 开源模型的人是一份难得的选型与验证清单。
Sources: julien.org | airealist.ai
Google Argon agent 自主优化数据中心内存,释放 300TiB | agent 自主运维生产基础设施的第一个量化样本
Google Gemini 4 Argon 博客披露:Argon agent 自主分析数据中心 profiling telemetry,跨多个 Google 数据中心应用内存优化,释放 300TiB 内存,预计总节省 500TiB–1PiB。这是「agent 自主优化生产基础设施」目前少见的具体数字样本——不是 demo,而是跑在真实 fleet 上的收益。对做 agent 落地与 infra 自动化的团队,可作为「agent 价值如何量化」的参照锚点。
Sources: registerspill.thorstenball.com
Anthropic 评测 GLM-5.3 网络攻击能力:4% 试验实现完整控制流劫持 | 开源模型攻击能力逼近闭源前沿的信号
Anthropic 对 GLM-5.3 的网络攻击能力做了评测:GLM-5.3 在 4% 试验中实现完整控制流劫持,接近 Claude Mythos Preview 的 6%,而更早的 Opus 4.6 与 GLM-5.2 成功率均为零;且 GLM-5.3 的防护在 64%–100% 的模拟测试中被简单技术绕过。对关注模型安全边界、红队评测与开源权重风险的从业者,这是一条把「开源模型攻击能力正在追上闭源」落到具体百分比的一手材料。
Sources: registerspill.thorstenball.com
AgentGuardBench:面向工具调用 agent 的多语言安全基准 | 安全基准不应奖励「一律拒绝」
作者发布 AgentGuardBench v0.1.1:120 条全合成场景,覆盖 prompt injection、隐私泄露、工具误用、权限越权、记忆安全与良性对照六类风险,横跨银行/医疗/教育/政府/招聘五个行业,支持英/法/斯瓦希里/约鲁巴四种语言,期望动作分为回答、脱敏、拒绝、请求审批四类。设计上强调「安全基准不应奖励一律拒绝」,因此保留良性任务完成度作为对照;所有工具为惰性、参考策略确定性,可完全本地运行、无需生产凭证或付费 API。输出机器可读的通过率、攻击成功率、隐私泄露率、未授权工具调用、人工审批违规、良性任务完成率等指标。适合关注 agent 安全评测与红队建设的团队作为自建评测集的参考模板。
Sources: dev.to
生成式推荐从零复现:Semantic ID + 生成式检索 + ranker 全管线 | 带 Colab 的 RecSys 工业范式 build log
Towards AI 上的一篇 RecSys 实战教程:用真实 Amazon 数据从零搭建生成式推荐管线——先做 popularity / ALS 基线,再引入 Semantic ID 给物品赋予语义身份,用小型 Transformer 做生成式检索(直接「写出」下一个 item 而非检索),最后叠加 ranker,并附数学推导与可交互 Colab notebook。对想理解 YouTube/Netflix/Meta 正在迁移的生成式推荐范式、并动手复现的从业者有直接参考价值。
Sources: pub.towardsai.net
Nvidia 盘中创新高,但折旧之争与表外负债成两条暗线 | H100 六年后残值或跌至 3 万美元
Nvidia 盘中创 237.88 美元新高,摩根士丹利恢复其为半导体首选股并给 300 美元目标价。文章聚焦两条被市场忽视的风险线:一是 AI 芯片折旧年限之争——Nvidia 称大客户已把 AI 服务器折旧从 4 年延至 6 年,Michael Burry 反驳称更高效芯片上市后旧硬件将暴跌,Barkr AI 数据显示 8 卡 H100 系统当前约值 32 万美元、6 年后跌至约 3 万美元;二是表外负债——Needham 分析师 Laura Martin 指出 Meta 表外承诺 6280 亿美元,并入 EV 后估值倍数抬升 35%,hyperscaler 表外承诺合计已达 3.8 万亿美元。适合关注 AI 资本开支可持续性的从业者拿到几个具体数字。
Sources: finance.biggo.com
Intel 年内股价翻倍:AI 算力需求外溢至通用 CPU 供应链 | CEO 称仅能满足约 50% CPU 客户需求
Intel 2026 年股价年内涨超 200%,Q2 营收 161 亿美元同比 +25%,数据中心与 AI 部门同比 +59% 至 63 亿美元,EPS 0.42 美元远超预期 0.21。CEO 陈立武称仅能满足约 50% 的 CPU 客户需求,年内已三次提价(含 10 月 5 日约 10% 的 PC 处理器涨价),Melius Research 将目标价上调至 165 美元。可作为「AI 算力需求外溢至通用 CPU 供应链」的一个量化注脚。
Sources: startupfortune.com
亚马逊再投 10 亿美元建 AI 数据中心,但一块 GPU 都不买 | 100+ 数据中心禁令下的社区关系投资
亚马逊宣布再投 10 亿美元用于 AI 数据中心,但资金全部投向职业培训、水资源保护和能源可负担性——一块 GPU 都不买。背景是全美 100 多个数据中心禁令正在被审议,社区对服务器农场的三大抱怨(就业、用水、电价)正是这笔钱的投向。AWS CEO Matt Garman 将数据中心建设类比州际公路系统。文章同时点出亚马逊与 Meta 上季度 AI 基建烧钱数百亿,但其中一家没有付费租户来回收账单。对关注数据中心选址阻力与 AI 基建公共关系的读者是一条具体信号。
Sources: 247wallst.com
🎙️ Podcast Picks
How to Choose Your Personal AI Agent
📍 Source: AI Daily Brief | ⭐ 3/5 | 🏷️ Agent, Product | ⏱️ 00:24:31
NLW compares personal AI agents like Dots, Muse, GrokBot, and OpenClaw across work vs personal use, model choice, deployment difficulty, and data privacy. There's also an interactive quiz to help you pick a starting point.
💡 Why Listen: Good if you're shopping for a personal agent and want a quick lay of the land. It's a consumer buying-guide angle though — light on technical depth.
📄 Paper Highlights
FinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms
JPMorgan Chase | 🏷️ Agentic Workflow, NLP Task, Inference
A hybrid LLM pipeline recovers missed trades from messy multi-party trading chatrooms, using fine-tuned classifiers as inference-time scaffolds and a difficulty-aware router to cut LLM calls by 85% — deployed at 70,000 RFQs/day.
🐙 GitHub Trending
No GitHub trending data provided today.