AI Tech Daily - 2026-07-17

Two massive open-source model launches reshaped the AI landscape today. Moonshot AI released Kimi K3, a 2.8T-parameter behemoth that tops Frontend Code Arena ahead of Claude Fable 5, while Thinking Machines Lab's Inkling (975B MoE) matches Nvidia's flagship at one-third the token cost. Meanwhile, Mi

AI Tech Daily - 2026-07-16

AI hit a major inflection point today: Thinking Machines Lab dropped Inkling, a 975B-parameter open-source MoE model, but early tests show it lags far behind Chinese frontier models and fails the Lem test — a basic reasoning benchmark every frontier model has passed since DeepSeek-R1. Meanwhile, Chi

AI Tech Daily - 2026-07-15

AI hit multiple milestones today. OpenAI's Codex hit 6M users (adding 1M daily), while GPT-5.6 sol slashed costs to a quarter of fable. Tencent open-sourced a 1-bit quantized 295B Hy3 model that runs on a single GPU with only 5% performance loss — Emad Mostaque called it the biggest news of the day.

AI Tech Daily - 2026-07-14

AI industry dynamics shifted fast today. Apple sued OpenAI for trade secret theft — Ben Thompson calls it a frustrated move masking Apple's deeper AI strategy problem. OpenAI GPT-5.6 Sol/Terra/Luna landed on Amazon Bedrock with big Agent benchmark gains. Microsoft dropped a 109-page MAI-Thinking-1 t

AI Tech Daily - 2026-07-13

AI's cost wars and safety debates dominated today's news. Li Auto's Mach-Mind-4-Flash proved a 35B MoE model (3B activated) can rival 100B-class models through post-training alone — a direct challenge to the scaling orthodoxy. Meanwhile, Oracle's S&P downgrade to BBB- (just above junk) was explicitl

AI Tech Daily - 2026-07-12

AI hit major milestones today: Anthropic's valuation surged past $1.2 trillion, overtaking OpenAI and kicking off its IPO — a seismic shift in the AI industry pecking order. Perplexity's CEO predicted model costs will drop 3-4x within 6-12 months, bringing Opus-level quality to local devices. Meanwh

AI Weekly 2026-W28

This week's core narrative is "release density meets engineering depth." OpenAI dropped GPT-5.6 as three models, ChatGPT Work, and GPT-Live — not a simple version bump, but a product matrix reorganization. Model capability tiers (Sol/Terra/Luna), Agent productization (Work), and interaction paradigm shift (full-duplex voice) all landed at once. Meanwhile, Agent engineering entered a "tool call refinement" phase: GitHub Copilot's postmortem, AWS's MCP design guide, Amazon and Writer's papers on orchestration efficiency — all point to the same judgment — an Agent's value no longer depends on whether it *can* call tools, but on *how well* it calls them. On inference acceleration, vLLM 0.25.0 runs 450+ Transformers architectures natively, DeepSeek's DSpark boosts generation speed by 60-85% under live traffic. These engineering deployments impact downstream decisions more than architecture papers.

RecSys Weekly 2026-W28

This week's recommendation system research centers on three technical threads: industrial deployment and theoretical deepening of generative retrieval, LLM/Agent moving from proof-of-concept to production, and robustness optimization of ranking/federated learning in industrial environments. Generative retrieval accelerates deployment with finer multi-interest modeling: Kuaishou deployed a heterogeneous generative architecture HGenPush in its push notification system, replacing traditional autoregressive decoding with non-autoregressive multi-token prediction, lifting DAU by 0.181%. Walmart introduced inventory-aware RAG into sponsored search, InvAwr-RAG boosting ad fill rate by 68%. On the theory side, BACH uses Bayesian mixture heads to solve the routing collapse problem in multi-interest two-tower models, achieving new recall SOTA on three benchmarks; DaV-Gen proposes a draft-and-verify mechanism unifying efficiency and accuracy in generative retrieval. Separately, Signed MaxSim is the first theoretical proof that MaxSim's expressiveness is at least as strong as vector inner products, and extends it to arbitrary real-valued inner products. LLM/Agent recommendations move from prototype to production: Meta's SCOReD is the week's most notable deployment — using student-aware CoT optimization to adapt teacher reasoning trajectories to small models, achieving +1.56% NDCG and +1.9% Recall@5 online while reducing reasoning length by 27.3%. Walmart used LLAMA2 7B + LoRA for three-category ad relevance classification, reaching 89.43% accuracy — surpassing GPT-4. Academically, MMEACR proposes a dual-track memory architecture to enhance agent visual reasoning; LBR systematically reveals length bias in LLM recommendations and offers a lightweight correction (NDCG@5 +16.82%); the survey Autonomous Information Seeking establishes a three-paradigm taxonomy for agent-based recommendation. Industrial ranking and federated learning optimization: Kuaishou's PIT-SUN is a deployable e

AI Tech Daily - 2026-07-11

AI hit a historic milestone today: OpenAI's GPT-5.6 Sol Ultra proved a 50-year-old unsolved math conjecture in under an hour using 64 parallel sub-agents — the first time a publicly available model has achieved a major mathematical breakthrough. Meanwhile, the agent infrastructure race intensified:

AI Tech Daily - 2026-07-10

Today marks a major inflection point in the AI industry. OpenAI dropped the GPT-5.6 family (Sol/Terra/Luna) alongside ChatGPT Work — a super app that directly challenges Anthropic's Claude Cowork. Meanwhile, SpaceXAI launched Grok 4.5, an Opus-class model purpose-built for coding and agent workflows

AI Tech Daily - 2026-07-09

AI voice interactions hit a turning point: OpenAI launched GPT-Live, a full-duplex speech model that listens and speaks simultaneously, with Sam Altman calling it "the magic of feeling human." NVIDIA's Nemotron topped the LangChain Deep Agents benchmark with 10x lower cost — all gains from engineeri

AI Tech Daily - 2026-07-08

AI's tectonic plates shifted today. Microsoft quietly replaced OpenAI and Anthropic models with its own in some apps — a strategic pivot that could reshape the API ecosystem. Meanwhile, Chinese AI models hit 30%+ US enterprise token share on OpenRouter, driven by DeepSeek and Z.ai's cost advantage.

1
...
34567
...
18