RecSys Weekly 2026-W29

This week's recommendation system research clusters around four technical themes: generative recommendation entering industrial deep waters, ranking models evolving toward long sequences and fine-grained semantics, retrieval systems breaking through on heterogeneous indexing and causal optimization, and LLM-enhanced recommendation moving from experiments to engineering deployment. Of the 34 papers, 23 come from industry (18 deployed), and 13 report online A/B results. Theme 1 "Generative Recommendation: From DocID Design to Fine-Tuning Alignment": Alibaba's CRID encodes business value ranking directly into DocIDs, achieving +1.06% GMV on a 300M item catalog at full traffic. GFlowGR fine-tunes generative recommendation with GFlowNet, delivering +0.4% annual revenue in Taobao search ads. Meituan's NONTP extends NTP training signals via temporal contrastive learning and cross-domain learning, lifting online CTR by +1.8% and GMV by +2.1%. Common thread: generative recommendation is shifting from "being able to generate" to "optimizing better." Theme 2 "Ranking Models Pursue Deep Decoupling and Long-Term Modeling": Meta's SlimPer formulates personalized ranking as iterative refinement of a <user, item> knowledge base, supporting 10k+ historical events with O(N) complexity, deployed on Instagram. Yandex's Long-History User Transformers decouple long-history inference via offline encoding + caching + a lightweight online model, achieving +2.77% in search ads. Alibaba's SAM uses satiety-gated explicit modeling of interest lifecycles, reducing post-purchase repetition rate by 60%. Theme 3 "Engineering and Causal Paradigms in Retrieval": Pinterest's causal retrieval framework reduces shopping triggers by 85% without harming key sessions. MESH uses modular architecture and gated bias correction to boost the scaling exponent for fresh items by 14x, with user retention +0.46%. Microsoft's FlashTrie fully migrates constrained decoding for generative retrieval to GPU, handling an

AI Tech Daily - 2026-07-18

AI economics is shifting fast. OpenAI proposed "Useful Intelligence per Dollar" as the new ROI metric, while NVIDIA countered with "intelligence per dollar" for post-training workloads. Anthropic is reportedly in talks to lease $10B in compute from Meta, and a $400M deal marks the first major GPU fi

AI Tech Daily - 2026-07-17

Two massive open-source model launches reshaped the AI landscape today. Moonshot AI released Kimi K3, a 2.8T-parameter behemoth that tops Frontend Code Arena ahead of Claude Fable 5, while Thinking Machines Lab's Inkling (975B MoE) matches Nvidia's flagship at one-third the token cost. Meanwhile, Mi

AI Tech Daily - 2026-07-16

AI hit a major inflection point today: Thinking Machines Lab dropped Inkling, a 975B-parameter open-source MoE model, but early tests show it lags far behind Chinese frontier models and fails the Lem test — a basic reasoning benchmark every frontier model has passed since DeepSeek-R1. Meanwhile, Chi

AI Tech Daily - 2026-07-15

AI hit multiple milestones today. OpenAI's Codex hit 6M users (adding 1M daily), while GPT-5.6 sol slashed costs to a quarter of fable. Tencent open-sourced a 1-bit quantized 295B Hy3 model that runs on a single GPU with only 5% performance loss — Emad Mostaque called it the biggest news of the day.

AI Tech Daily - 2026-07-14

AI industry dynamics shifted fast today. Apple sued OpenAI for trade secret theft — Ben Thompson calls it a frustrated move masking Apple's deeper AI strategy problem. OpenAI GPT-5.6 Sol/Terra/Luna landed on Amazon Bedrock with big Agent benchmark gains. Microsoft dropped a 109-page MAI-Thinking-1 t

AI Tech Daily - 2026-07-13

AI's cost wars and safety debates dominated today's news. Li Auto's Mach-Mind-4-Flash proved a 35B MoE model (3B activated) can rival 100B-class models through post-training alone — a direct challenge to the scaling orthodoxy. Meanwhile, Oracle's S&P downgrade to BBB- (just above junk) was explicitl

AI Tech Daily - 2026-07-12

AI hit major milestones today: Anthropic's valuation surged past $1.2 trillion, overtaking OpenAI and kicking off its IPO — a seismic shift in the AI industry pecking order. Perplexity's CEO predicted model costs will drop 3-4x within 6-12 months, bringing Opus-level quality to local devices. Meanwh

AI Weekly 2026-W28

This week's core narrative is "release density meets engineering depth." OpenAI dropped GPT-5.6 as three models, ChatGPT Work, and GPT-Live — not a simple version bump, but a product matrix reorganization. Model capability tiers (Sol/Terra/Luna), Agent productization (Work), and interaction paradigm shift (full-duplex voice) all landed at once. Meanwhile, Agent engineering entered a "tool call refinement" phase: GitHub Copilot's postmortem, AWS's MCP design guide, Amazon and Writer's papers on orchestration efficiency — all point to the same judgment — an Agent's value no longer depends on whether it *can* call tools, but on *how well* it calls them. On inference acceleration, vLLM 0.25.0 runs 450+ Transformers architectures natively, DeepSeek's DSpark boosts generation speed by 60-85% under live traffic. These engineering deployments impact downstream decisions more than architecture papers.

RecSys Weekly 2026-W28

This week's recommendation system research centers on three technical threads: industrial deployment and theoretical deepening of generative retrieval, LLM/Agent moving from proof-of-concept to production, and robustness optimization of ranking/federated learning in industrial environments. Generative retrieval accelerates deployment with finer multi-interest modeling: Kuaishou deployed a heterogeneous generative architecture HGenPush in its push notification system, replacing traditional autoregressive decoding with non-autoregressive multi-token prediction, lifting DAU by 0.181%. Walmart introduced inventory-aware RAG into sponsored search, InvAwr-RAG boosting ad fill rate by 68%. On the theory side, BACH uses Bayesian mixture heads to solve the routing collapse problem in multi-interest two-tower models, achieving new recall SOTA on three benchmarks; DaV-Gen proposes a draft-and-verify mechanism unifying efficiency and accuracy in generative retrieval. Separately, Signed MaxSim is the first theoretical proof that MaxSim's expressiveness is at least as strong as vector inner products, and extends it to arbitrary real-valued inner products. LLM/Agent recommendations move from prototype to production: Meta's SCOReD is the week's most notable deployment — using student-aware CoT optimization to adapt teacher reasoning trajectories to small models, achieving +1.56% NDCG and +1.9% Recall@5 online while reducing reasoning length by 27.3%. Walmart used LLAMA2 7B + LoRA for three-category ad relevance classification, reaching 89.43% accuracy — surpassing GPT-4. Academically, MMEACR proposes a dual-track memory architecture to enhance agent visual reasoning; LBR systematically reveals length bias in LLM recommendations and offers a lightweight correction (NDCG@5 +16.82%); the survey Autonomous Information Seeking establishes a three-paradigm taxonomy for agent-based recommendation. Industrial ranking and federated learning optimization: Kuaishou's PIT-SUN is a deployable e

AI Tech Daily - 2026-07-11

AI hit a historic milestone today: OpenAI's GPT-5.6 Sol Ultra proved a 50-year-old unsolved math conjecture in under an hour using 64 parallel sub-agents — the first time a publicly available model has achieved a major mathematical breakthrough. Meanwhile, the agent infrastructure race intensified:

AI Tech Daily - 2026-07-10

Today marks a major inflection point in the AI industry. OpenAI dropped the GPT-5.6 family (Sol/Terra/Luna) alongside ChatGPT Work — a super app that directly challenges Anthropic's Claude Cowork. Meanwhile, SpaceXAI launched Grok 4.5, an Opus-class model purpose-built for coding and agent workflows

1
...
89101112
...
24