RecSys Weekly 2026-W38
2026-9-19
| 2026-9-19
字数 4996阅读时长≈ 13 分钟
type
Post
status
Published
date
Sep 19, 2026 05:31
slug
rec-weekly-en-2026-W38
summary
This week's recommender systems research clusters around three threads: industry migrating generative paradigms into existing ranking/retrieval pipelines, user semantic signals (natural-language rationales, reasoning traces) entering the recommendation feature space, and long-sequence compression with self-evolving memory. Thread one: generative upgrades take the "smooth migration" path. Meta's LIGE-GR generalizes a mature pointwise ranking system into listwise generation and evaluation, lifting Instagram Reels time spent by 1.14% and Facebook Video by 0.72% — with only modest extra inference cost. Alibaba International Digital Commerce's LazFormer uses generative pretraining to provide both sparse and dense parameter initialization for ranking, then adds a transferable residual adapter to fix dense-parameter negative transfer. Both point the same way: don't replace the serving infrastructure, just rewrite the representation and the optimization target. Thread two: user semantic signals become first-class text signals. Kuaishou's SARA scales user natural-language preference rationales (AURs) from 86,564 authors to a 10M-author space, and feeds positive and negative rationales into production ranking. Alibaba's CoFree targets reasoning collapse in LLM embeddings, coupling embedding-oriented and reasoning-oriented dual rewards end-to-end. CoFree-4B gains an average absolute +2.8 nDCG@10 over Qwen3-Embedding-4B across 22 datasets spanning MTEB and BRIGHT. Thread three: long-sequence compression and memory isolation. Tencent's ChronicleRec compresses ultra-long behavior sequences into time-anchored Chronicle Tokens in one pass, caches them per user, and decouples ultra-long sequence modeling from online candidate scoring. LION names evolution conflict — heterogeneous preference drift fighting inside a shared autoregressive parameter space, where dominant behavior patterns suppress long-tail ones. The fix is parameter isolation via a sparse Key-Value memory layer.
tags
Recommendation Systems
Weekly
Papers
category
Rec Tech Report
icon
📚
password
priority
1

Weekly Overview

This week's recommender systems research clusters around three threads: industry migrating generative paradigms into existing ranking/retrieval pipelines, user semantic signals (natural-language rationales, reasoning traces) entering the recommendation feature space, and long-sequence compression with self-evolving memory.
Thread one: generative upgrades take the "smooth migration" path. Meta's LIGE-GR generalizes a mature pointwise ranking system into listwise generation and evaluation, lifting Instagram Reels time spent by 1.14% and Facebook Video by 0.72% — with only modest extra inference cost. Alibaba International Digital Commerce's LazFormer uses generative pretraining to provide both sparse and dense parameter initialization for ranking, then adds a transferable residual adapter to fix dense-parameter negative transfer. Both point the same way: don't replace the serving infrastructure, just rewrite the representation and the optimization target.
Thread two: user semantic signals become first-class text signals. Kuaishou's SARA scales user natural-language preference rationales (AURs) from 86,564 authors to a 10M-author space, and feeds positive and negative rationales into production ranking. Alibaba's CoFree targets reasoning collapse in LLM embeddings, coupling embedding-oriented and reasoning-oriented dual rewards end-to-end. CoFree-4B gains an average absolute +2.8 nDCG@10 over Qwen3-Embedding-4B across 22 datasets spanning MTEB and BRIGHT.
Thread three: long-sequence compression and memory isolation. Tencent's ChronicleRec compresses ultra-long behavior sequences into time-anchored Chronicle Tokens in one pass, caches them per user, and decouples ultra-long sequence modeling from online candidate scoring. LION names evolution conflict — heterogeneous preference drift fighting inside a shared autoregressive parameter space, where dominant behavior patterns suppress long-tail ones. The fix is parameter isolation via a sparse Key-Value memory layer.

Generative Recommendation and Ranking Paradigm Upgrades

This chapter covers different facets of one question: how the generative paradigm enters an existing industrial recommendation stack, rather than replacing it.
SARA (Kuaishou) — treats Articulated User Rationales (AURs) as polarity-aware, reason-level text signals. Three stages. A data engine collects and filters AURs from 240M Kuaishou Live users, producing the author-centric, quality-controlled SARA-HQ dataset. A general MLLM is aligned into SARA-7B via large-scale SFT plus Quality-Refining DPO, expanding rationale coverage from 86,564 AUR-covered authors to the full 10M-author space. SARA-Ranker then wires the generated positive and negative rationales into production ranking through rationale-aware interaction modeling and rejection-memory modeling. It has been deployed online for over 30 days with daily refreshes; A/B shows higher engagement and lower negative feedback. This line continues S²GR, which used thinking tokens for semantic guidance on Kuaishou short video — but swaps the supervision signal from a semantic codebook to reasons users wrote themselves. The related EASQ uses survey signals for satisfaction alignment; SARA can be read as pushing survey sparsity further, out to the full author space. Worth noting on the engineering side: rationales are naturally sparse, uneven in quality, and cover few items. SARA's value is proving that a "low-coverage but high-semantic-density" signal can be amplified by an MLLM and fed back into ranking.
LIGE-GR (Meta) — generalizes an existing pointwise ranking system into a listwise generation and evaluation system without rebuilding the stack. The paper names two real constraints: how sequence-level generation and optimization fit into recommendation, and the organizational and technical risk of wholesale replacement for a mature system that has iterated for years around product, business constraints, serving infrastructure, and team ownership. LIGE-GR keeps compatibility with the existing model, value function, and serving infrastructure, and only does generation and evaluation at the list level. Instagram Reels time spent +1.14%, Facebook Video +0.72%. The same "migrate, don't rewrite" line includes RelayGR, which used cross-stage relay inference to raise long-sequence length 1.5× and throughput 3.6×, and GRLM, which used structured Term IDs to unlock native recommendation potential. LIGE-GR's significance isn't algorithmic novelty — it's a replicable, minimally invasive upgrade template.
LazFormer (Alibaba) — a unified generative pretraining plus ranking framework, designed against three pain points in existing pretraining schemes. First, a transferable residual adapter injects ranking-specific features residually, easing dense-parameter negative transfer caused by mismatched pretraining and ranking input features. Second, a request-aware ranking module integrates long-sequence compression, hybrid sparse attention, and a request-aware paradigm to model long user sequences efficiently. Third, an asymmetric multi-round training strategy resets sparse parameters between rounds while continuously accumulating dense parameters, easing sparse-parameter overfitting. It's deployed in Alibaba International Digital Commerce's industrial recommender with online gains. Historically, DSIN used session partitioning plus Bi-LSTM to model interest evolution, and RAT used a retrieval-augmented Transformer to cover long-tail features. LazFormer's difference is treating parameter initialization itself as the transfer object.
GESE (Baidu) — headline generation for a commercial platform with 100M+ DAU, decoupling personalized title generation into "generation for exploration" and "selection for exploitation." The generation side treats the LLM as a probabilistic explorer, using Group Sequence Policy Optimization (GSPO) with hierarchical rewards to produce a candidate set that maximizes semantic coverage of latent user interest. The selection side uses a lightweight, real-time feedback-aware selector to pick the best realization for the current context from the candidate pool. Directly optimizing a single title causes mode collapse — convergence to a generic pattern that satisfies average taste, suppressing long-tail audience needs. Online CTR +2.57%, dwell time +0.87%. This shares its thinking with HyMiRec, whose two-tier structure uses a lightweight recommender and an LLM recommender to capture coarse and fine-grained interest separately: split "diversity exploration" and "precision exploitation" into two independently optimizable modules.
Verifier — a posterior verifier for first-stage retrieval in multi-stage systems. The problem is cleanly defined: the actual consumed prefix of a first-stage candidate list is bounded by the compute of downstream, more expensive ranking stages, so relevant items may sit deeper in the retrieval list but never enter the consumed prefix. The approach freezes the retriever as a drafter and trains a lightweight generative verifier that scores candidates by the likelihood of candidate-identifier tokens, running inference only on top-K candidates. Training uses next-token cross-entropy — no negative sampling or candidate pool — and it's compatible with any fixed tokenization. On Amazon and YaMBDa, one training recipe uniformly improves Recall@10 across four retrievers: SASRec, GRU4Rec, NextItNet, and MiniOneRec. Ablations show the gain isn't just from injecting content features into the retriever. The minimal interface is the engineering selling point — the retriever only supplies a query state and candidate items, and is never retrained.
Preference-Drift — efficiency and noise in long-sequence generative recommendation. Two empirical observations: full-sequence modeling keeps rising in compute cost with length, while accuracy gains saturate fast or even degrade from noise; target-aware context retrieval shortens input but is easily disturbed by "semantically consistent but preference-inconsistent" noise and incomplete context. The method learns soft subsequence boundaries from multi-dimensional preference drift information, aggregates items within a subsequence into preference-consistent representations using linear attention with soft assignment weights — avoiding full-sequence attention cost — then uses cross-attention to capture dependencies between recent interactions and relevant subsequence context, and finally gates the fusion of recent and global subsequence context. It complements xGR, which does staged computation and separated KV caches on the serving side — one changes serving scheduling, the other changes input organization.
LION — the first to name evolution conflict in generative recommendation: heterogeneous users' preference drift is optimized together inside a fully shared autoregressive parameter space, so dominant behavior patterns gradually dominate model evolution while under-represented patterns are persistently ignored. The framework centers on a sparse Key-Value memory layer, using sparse memory activation to isolate the evolution of different behavior patterns, plus a consolidation loss to strengthen learning of under-represented preference dynamics during continual adaptation. The paper lays out three design principles: isolated memorization, reinforced evolution, scalable application. Experiments on real datasets cover per-period evaluation, user/item group evaluation, and evolution convergence analysis. Where RelayGR tackles long-sequence inference latency and OneLoc tackles geo-aware semantic IDs, LION handles parameter competition along the time dimension — a previously overlooked class of problem.

Generative Retrieval and Recall Optimization

The core disagreement in this chapter: should semantic IDs be generated by an LLM, or replaced by the LLM's text representations?
VARG (Alibaba) — generative retrieval for Tmall App search, feeding generated item candidates straight into the existing fine-ranking stage. VARG-ID builds semantic prefixes with RQ-VAE, strengthens search relevance via bidirectional query-item contrastive learning, then appends a value-ordered third token that carries both fine-grained item address and business-value priors. Training has three stages: three-stage SFT learns item-to-identifier mapping, query semantic retrieval, and personalized retrieval in turn; personalized training combines value-aware and hierarchical alignment supervision, using LO-SFT (local ordinal supervision) to learn the intra-cluster local order encoded in the third token; Prefix-GRPO uses a gated reward composed of output validity, user behavior, ranker advantage, and search relevance, with prefix-aware token weighting, aligning candidate generation to business value and ranking objectives. Collaborative daily updates incorporate new items and behavior feedback while preserving existing item addresses. A 14-day, 20%-traffic A/B: GMV +1.45%, per-user IPV +0.22%, PCTR +0.31%. This line follows DualGR's dual-branch long/short-term interest and search-based SID decoding, but extends the third token from a pure semantic address into a carrier of value priors.
ANGLE (Tencent) — real-time retrieval for sponsored search ads, going the opposite direction from VARG: replacing discrete SIDs with LLM-generated hierarchical text representations. The hierarchy has two levels: commercial intent (high-level overview) and ad abstract (fine-grained detail). The paper's critique of the SID route is specific — SIDs aren't learned by the base LLM, the SFT stage must memorize a large SID-to-ad mapping, generalization to unseen ads is poor, maintenance and update costs are high, the one-to-one SID-to-ad mapping makes decoding inefficient, and reliance on a small reward model (e.g. pCTR) to assess relevance limits the LLM's full judgment of commercial value. ANGLE integrates retrieval, relevance, and ranking into a single LLM. Online spend +1.81%, GMV +2.16%; offline it outperforms 7 baselines on HR, ACR, and other metrics. Compared with Gryphon, which jointly trains SID generation and item-level scoring to ease "likelihood vs. relevance objective mismatch," ANGLE simply abandons the SID abstraction layer.
MIMA (Alibaba) — interest collapse in multi-interest retrieval. The paper fingers the single-positive paradigm as a major cause: each instance provides only one positive, interests are optimized independently, the same best-matching interest gets repeatedly updated toward different positives while the rest get insufficient supervision; and existing methods rarely model each user's activation strength per interest, so cross-interest channel routing scores aren't comparable at inference. MIMA encodes co-occurring items within the same request into a positive set, uses a causal Transformer decoder to generate complementary interests, and assigns each positive exclusively to a distinct interest via Hungarian matching — so interest differentiation comes from the training objective itself rather than auxiliary regularization. A lightweight routing module then estimates user-interest activation probabilities to calibrate cross-channel scores. It outperforms MIND, ComiRec, SINE, and PIMIRec on three public datasets plus one industrial dataset, with online A/B business gains. This shares its lineage with SetMIR, which models multi-interest retrieval as set prediction, trains with Hungarian matching, and dynamically decides the number of ANN queries.
CoFree (Alibaba) — targets reasoning collapse in LLM embedding learning, where specialization toward the embedding objective suppresses effective reasoning generation or produces retrieval-irrelevant text. Two stages: the first is reference-guided SFT, restoring the base embedding model's reasoning ability while preserving representation strength; the second introduces embedding-oriented and reasoning-oriented dual rewards, ensuring fine-grained reasoning toward the relevance objective during RL. This endpoint-coupled optimization turns embedding learning from static alignment into a high-quality, reasoning-guided retrieval search process. CoFree-4B gains an average absolute +2.8 nDCG@10 across 22 datasets spanning MTEB and BRIGHT, with consistent gains in online experiments on a real retrieval system. Neighboring work ReinPool uses RL to compress multi-vector embeddings into a single vector at 746–1249× compression, recovering 76–81% of multi-vector retrieval performance. Both answer the same question: how should the embedding training objective align with multi-stage downstream use?
InitGen (OPPO) — candidate generation for the interaction-initiation stage of an intelligent assistant. The core difficulty is incomplete feedback: the generator produces more candidates than are ultimately shown, and after downstream filtering and ranking only a subset reaches users, so observed feedback is partial and can't be reliably attributed to a single query. InitGen jointly generates a candidate query set and aligns to user feedback with weighted preference optimization, where weights come from user activity and downstream ranking scores — activity weights reduce the dominance of highly active users in training, and ranking scores serve as a practical estimate of feedback reliability. A rolling-window update folds recent interaction data into periodic model updates. Online A/B against a strong production baseline: CTR from 0.95% to 1.61% (relative +69.1%), query impressions +17.9% under the same traffic allocation, a full candidate set generated within 180 ms, serving 150M+ MAU. That CTR doubling is the largest relative gain among this week's industrial papers, but the base is low (0.95%) and needs to be read in context. Related prior work DualGR handles a similar problem with an exposure-aware next-token prediction loss.
PCap (Meta) — moves personalized diversity constraints from the ranking stage forward to retrieval. PCap models user-level diversity preference with Shannon entropy and buckets users, applying personalized category caps during multi-source candidate retrieval. For the high-dimensional parameter space formed by per-bucket caps, it navigates with an automated online optimization method, Parameter Tuning Sequence. Large-scale A/B on Facebook Marketplace shows statistically significant gains in engagement metrics. Where Disaggregated Multi-Tower trades topology-aware modeling for large-scale efficiency, PCap's contribution is introducing a user-granular "quota" constraint form at the retrieval stage.
ReWAM — retrieval deployment for general multimodal embeddings. The paper identifies two obstacles to corpus-scale deployment in explicit-CoT UME methods: GRPO gives all CoT tokens the same advantage, so it can't distinguish which tokens are input-supported evidence that separates positives from hard negatives; and generating a full CoT before every embedding adds substantial latency, even when partial traces already provide enough retrieval evidence. ReWAM uses Retrieval-aware Self-Distillation to construct privileged guidance from input-supported evidence, with an on-policy self-teacher refining trajectory-level feedback into token-level supervision. Retrieval-adaptive Inference uses a retrieval confidence head to estimate the remaining retrieval utility of a partial CoT, stopping unproductive traces early, and uses speculative decoding to accelerate productive continuations. It reaches SOTA on MMEB-V2 and MRMR, with inference throughput up to 5× that of explicit-CoT UME methods. Related work includes Kuaishou's DAS, which uses dual-aligned semantic IDs to optimize quantization and alignment.

Fine-Ranking Systems and Ad/Commerce Optimization

This chapter shares a full-page, full-funnel perspective — no longer optimizing a single item's score in isolation.
PinDCO (Pinterest) — a production-grade dynamic creative optimization system. Generative AI has sharply accelerated ad creative production, and the number of candidate variants per campaign has exploded, imposing matching requirements on DCO systems under latency and cost constraints. The core is the Creative Component Fusion Network (CCFN): each creative component (image, headline, layout) is modeled by its own tower, with component-level hyperparameters reflecting differing modeling complexity; fused component representations predict a creative-level score conditioned on the ad-level prediction. Training data quality is improved via an exploration-exploitation strategy. For Pinterest's waterfall grid layout — where creative render size affects neighboring content and session-level engagement — a Pixel-aware Adjustment Module adjusts scores by creative size to encourage efficient use of screen space and better whole-page outcomes. Supporting the large candidate volume are a lightweight pre-selection model for early pruning, plus caching and dynamic batching. Online A/B: ad CTR +3.09%, whole-page metrics positive, now live in Pinterest Ads.
UVR (Wolt) — merchant ranking in on-demand delivery. Unlike pure digital domains, candidate merchants are localized and constrained by real-time availability and delivery operations. The core tension is giving new merchants exposure to probe while preserving ranking quality in sessions with repurchase intent. The Universal Venue Ranker pairs a bidirectional Transformer sequence encoder with a GBDT ranker integrating context, user, and merchant features; training covers all merchants and domains in a country, and inference applies local delivery constraints. Label smoothing and trial-biased sample weighting push the model toward new merchants. Offline trial MRR improves +12% to +30% relative to production, at the cost of reorder MRR regression in five of six countries — but Global CVR (the core online metric mixing trial and reorder sessions) is statistically unchanged. Three consecutive A/Bs: V1 delivered +5.5% Merchant Trial Rate and +0.16% Global CVR, V2 added +0.45% MTR, and V3 unified food and retail rankers across domains for another +1.31% Retail MTR while replacing four separate rankers (three food, one retail). The tradeoff is worth recording: partial offline regression was accepted because the online blended metric held flat — a textbook case of system-level objective substitution.
ChronicleRec (Tencent) — a pretrain-transfer framework for lifelong user modeling. Existing lifelong-interest methods retrieve target-relevant behavior per candidate, coupling long-sequence modeling with candidate scoring and duplicating online cost. Recent target-agnostic compression methods support caching user summaries, but often append a query token at the sequence end and encode bidirectionally, producing unordered, redundant summaries that ignore temporal structure. ChronicleRec compresses ultra-long behavior sequences in one pass into time-ordered Chronicle Tokens: recency-aware multi-granularity merging preserves recent behavior while coarsening distant history; query tokens are interleaved with the merged sequence, and a causal encoder lets each query summarize only history before its time anchor; a multi-horizon design masks different recent-history windows across parallel branches to learn complementary long-range interests; and the compressor is pretrained with a mask-and-predict objective, reconstructing held-out recent behavior from compressed old history to align historical signals with near-term intent. Because Chronicle Tokens are target-agnostic, they can be cached per user, decoupling ultra-long sequence modeling from online candidate scoring. It outperforms recent-window and single-pass compression baselines on KuaiRand and Tencent AdLive, approaching full-attention performance, with a seven-day online A/B confirming gains. This shares its question with CASE, which separates item-level rhythm learning from cross-item interaction: how do you compress long-term history without losing temporal structure?
Single-Token (Indeed) — an ordinal scoring primitive for cold-start candidate ranking. The scenario constraints are specific: a recruiting sourcing platform has low, niche traffic and can't produce the millions of logged interactions a traditional deep neural ranker needs; zero-shot LLM scores are unstable, non-deterministic, and insufficiently accurate for ranking; what's available is a few hundred thousand ordinal relevance labels. The method models candidate-job relevance as ordinal classification over rating tokens {1,...,5}, with the relevance score as the expectation of the first token's probability distribution. Because the score comes from a single decoding step rather than open-ended generation, it's a deterministic function of model logits — no output parsing, low latency. To learn nonlinear interdependencies of heterogeneous hiring standards from this supervision, it fine-tunes a small language model with a hybrid ordinal regression loss combining an MSE term (preserving ordinal distance) and a categorical cross-entropy term (sharpening class boundaries), evaluated along Jobseeker Relevance and Employer Relevance with NDCG@10 and low-relevance rate. Offline it outperforms heuristic baselines and zero-shot LLMs; end-to-end simulation shows the same direction with larger magnitude (Jobseeker NDCG@10 +54.2%, low-relevance rate -46.7%); online experiments cut employer low-relevance by 27.3% and lift keep rate by 7.07%. The combination of ordinal regression and single-step decoding is uncommon on the recommendation side, and directly relevant to low-traffic cold-start scenarios.
Query Suggestion (Alibaba) — two-stage optimization for generative query suggestions. The central challenge: the generated slate must both make each query useful and cover diverse intents. Stage one is intent-aware diversity modeling: construct intent-aligned SFT data and optimize intent coverage with an Intent-Aware Diversity Reward, modeling diversity in semantic intent space rather than at the text level. Stage two is query-level credit assignment: route individual quality signals to the corresponding query tokens, while the slate-level diversity signal is shared across the whole slate. Offline evaluation on a large-scale production dataset and online A/B both show gains in CTR, query quality, and intent coverage. Where AdaGRPO does adaptive loss gating to handle reward noise in generative recommendation, the focus here is reward allocation granularity rather than magnitude.

LLM Agents and Retrieval-Augmented Generation

All three papers in this chapter swap fixed pipelines for adaptive decision-making. The common element is an "intermediate decision layer."
AURA — an end-to-end agentic diagnosis and code-level improvement system for production recommenders. The motivation is blunt: aggregate metrics (AUC, NDCG, precision/recall, diversity, coverage) give a high-level and incomplete picture, and seeing where and for which users a recommender fails requires reasoning at scale with domain understanding. AURA's specialized agents read production engagement logs (from thousands to millions of sessions), surfacing patterns and examples of real users where recommendations failed; next, those diagnoses plus context on the recommender's own code, data, and training pipeline are used to propose and implement improvements grounded in the codebase. The paper reports system design, preliminary testing on production data from two large consumer platforms at a major media streaming company, safeguards, and operational experience. The authors stress the architecture is transferable — all domain-specific elements are injected through a configuration layer, already ported between the two platforms, with a mapping given to e-commerce and online retail. Earlier, DualAgent-Rec used an LLM as coordinator to schedule development/exploration dual agents for multi-objective constrained optimization, and RES used a Reasoner-Executor-Synthesizer three-layer architecture to separate intent parsing, deterministic retrieval, and narrative generation to eliminate data hallucination. AURA goes one step further: after diagnosis, it edits code directly. The judgment to preserve: this is a workshop paper reporting early results rather than a full A/B, and it lacks quantified gains.
MemRetriever — models long-term memory access as a multi-step search process. The paper critiques three failure modes of static top-k memory retrieval: a single query, a fixed return count, and direct handoff downstream — which miss evidence scattered across sessions, introduce irrelevant content, and waste context, especially for multi-hop, temporal, and knowledge-updating questions. At each step, MemRetriever reasons over existing evidence and chooses parallel search for breadth exploration, serial search for targeted completion, or reflection and denoising for filtering and evidence assessment, stopping once retained evidence suffices to answer downstream. Training uses ReAct-style search-memory traces for supervised warm-start, then GRPO optimization, with rewards encouraging evidence coverage, noise reduction, answer sufficiency, and efficient termination. It consistently outperforms static retrieval and supervised-only baselines on LOCOMO, LongMemEval, HotpotQA, MuSiQue, and 2WikiMultiHopQA; MemRetriever-4B-RL exceeds DeepSeek-v4-Flash on LongMemEval's main retrieval metric under the same pipeline. The decision logic is storage-backend-agnostic and can run on top of external knowledge bases and vector databases. Related LLM recommendation alignment work includes ITPO, which derives turn-level process rewards from sparse outcome signals, and WPAUC, which proves GRPO optimizing an LLM recommender is equivalent to maximizing AUC and improves the objective accordingly.
Pre-retrieval Query Clustering (Amazon) — query-adaptive retrieval depth for RAG systems. The fragility of fixed top-k: simple queries over-retrieve (introducing noise and cost), complex queries under-retrieve (recall failures cascade into wrong answers). The method splits into offline and online stages. Offline, it estimates per-query retrieval difficulty, measuring NDCG under the default retriever and deriving a query-specific "saturation point" k* from the NDCG-k curve. Because computing these signals online is expensive, it clusters large-scale queries in embedding space and uses a mean-plus-variance rule to summarize, per cluster, a recommended retrieval depth targeting high coverage (~95%). At runtime, an incoming query is assigned to a cluster and the corresponding top-k is selected in constant time. Compared with posterior confidence methods that rely on already-retrieved documents for clustering, this is a pre-retrieval, query-centric approach — more robust on heterogeneous, case-based corpora, and transferable to legal, medical, financial, and enterprise search. In full-traffic query tests, F1 improves by over 36%, and token usage drops 14% in low-complexity clusters with no accuracy loss. The same "adaptive stopping" motif appears in ReRec and Rank-GRPO as difficulty curriculum scheduling and rank-unit rewards, respectively.

Directions to Watch

Will "minimally invasive upgrades" for generative ranking become a standard engineering paradigm? This week, Meta's LIGE-GR, Alibaba's LazFormer, and Tencent's ChronicleRec all chose the same strategy — keep the existing serving infrastructure, replace only the representation layer or output organization, rather than rebuilding the stack. This converges with RelayGR's staged inference optimization on the serving side and GRLM's use of structured Term IDs to unlock native recommendation potential. All three teams report online gains in the 0.7%–3% range — modest, but risk-controlled. If this "smooth migration" keeps producing stable small positive returns, industry adoption of generative recommendation may diverge from the DLRM replacement path — closer to incremental refactoring.
User semantic signals and reasoning signals are splitting into two technical sub-lines. One is explicit rationales: SARA uses rationales across a 10M-author space and rejection-memory modeling for negative feedback, with related work in EASQ's survey-based satisfaction alignment and TagLLM's fine-grained label generation. The other is reasoning traces: CoFree uses dual rewards to preserve embedding reasoning quality, and ReWAM uses retrieval feedback for token-level credit assignment with adaptive early CoT termination. Both sub-lines must answer the same engineering question: how do you price the extra semantic/reasoning compute? ReWAM's 5× inference throughput and CoFree's +2.8 nDCG@10 are the clearest pricing references so far.
Where sparse memory and parameter isolation fit in continual learning. LION's evolution conflict is a problem that hadn't been explicitly named before — generative recommendation's autoregressive parameter space is inherently shared, and heterogeneous preference drift competes within it. A sparse Key-Value memory layer refines isolation granularity from "task" down to "behavior pattern," paired with a consolidation loss to strengthen under-represented dynamics. This direction shares a design philosophy with SetMIR, which uses set prediction plus Hungarian matching to force interest differentiation, and MIMA, which drives interest differentiation through exclusive assignment: let the supervision objective itself produce structural separation. For scenarios with continual evolution needs and heavy long tails (live streaming, short video, local life), this line's deployment potential is worth tracking.

Paper Roundup

Generative Recommendation and Ranking Paradigm Upgrades
SARA — Kuaishou builds a 240M-user data engine to produce SARA-HQ, aligns SARA-7B via SFT + Quality-Refining DPO to expand rationale coverage from 86,564 authors to 10M authors; SARA-Ranker lifts engagement and cuts negative feedback in production ranking, deployed with daily refreshes for over 30 days.
LIGE-GR — Meta generalizes mature pointwise ranking into listwise generation and evaluation; Instagram Reels time spent +1.14%, Facebook Video +0.72%, with only modest extra inference cost.
LazFormer — Alibaba International Digital Commerce proposes a unified generative pretraining + ranking framework: a transferable residual adapter fixes negative transfer, a request-aware ranking module models long sequences, and asymmetric multi-round training eases sparse-parameter overfitting; deployed with online gains.
GESE — Baidu decouples exploration and exploitation in title generation on a 100M+ DAU platform: LLM + GSPO + hierarchical rewards for semantic coverage, a real-time selector for precision exploitation; CTR +2.57%, dwell time +0.87%.
Verifier — proposes a frozen retriever plus lightweight generative verifier, scoring top-K by identifier-token likelihood with no negative sampling; uniformly improves Recall@10 for SASRec, GRU4Rec, NextItNet, and MiniOneRec on Amazon and YaMBDa.
Preference-Drift — learns soft subsequence boundaries from multi-dimensional preference drift, aggregates preference-consistent representations with linear attention, and combines recent and long-term preference via cross-attention plus gated fusion; improves both accuracy and efficiency in long-sequence generative recommendation.
LION — first to name evolution conflict in generative recommendation; uses a sparse Key-Value memory layer for parameter isolation and a consolidation loss to strengthen under-represented preference dynamics, validated across multiple datasets for continual evolution.
Generative Retrieval and Recall Optimization
VARG — Alibaba Tmall search builds VARG-ID from RQ-VAE semantic prefixes plus a value-ordered third token, with three-stage SFT (including LO-SFT) + Prefix-GRPO gated rewards; in a 14-day, 20%-traffic A/B, GMV +1.45%, IPV +0.22%, PCTR +0.31%.
ANGLE — Tencent replaces discrete SIDs with LLM-generated hierarchical text representations (commercial intent + ad abstract), unifying retrieval/relevance/ranking in a single LLM; online spend +1.81%, GMV +2.16%.
MIMA — Alibaba International Digital Commerce turns multi-interest learning into exclusive multi-positive assignment, with Hungarian matching driving interest differentiation and a lightweight router calibrating cross-channel scores; beats SOTA on three public datasets plus an industrial dataset, with online A/B business gains.
CoFree — Alibaba solves reasoning collapse with reference-guided SFT plus embedding/reasoning dual-reward RL; CoFree-4B gains an average +2.8 nDCG@10 across 22 datasets spanning MTEB and BRIGHT, with consistent online gains on a real retrieval system.
InitGen — OPPO's Xiaobu assistant generates interaction-initiation candidates with weighted preference optimization (activity + downstream ranking scores); online CTR 0.95%→1.61% (relative +69.1%), impressions +17.9%, completed within 180 ms, serving 150M+ MAU.
PCap — Meta applies personalized category caps at the retrieval stage on Facebook Marketplace, modeling diversity preference with Shannon entropy and bucketing users, tuning online via Parameter Tuning Sequence; large-scale A/B shows significant engagement gains.
ReWAM — uses Retrieval-aware Self-Distillation to refine GRPO trajectory-level feedback into token-level supervision, and Retrieval-adaptive Inference to terminate inefficient CoT early with speculative decoding acceleration; SOTA on MMEB-V2 and MRMR, up to 5× inference throughput.
Fine-Ranking Systems and Ad/Commerce Optimization
PinDCO — Pinterest's production-grade DCO system: Creative Component Fusion Network for component-level modeling, Pixel-aware Adjustment Module for whole-page scoring in the waterfall grid; online CTR +3.09%, live in Pinterest Ads.
UVR — Wolt replaces four separate models with a bidirectional Transformer + GBDT hybrid ranker, using trial-biased sample weighting to balance probing and repurchase; three A/Bs cumulatively deliver +5.5% Merchant Trial Rate, +0.16% Global CVR, +1.31% Retail MTR.
ChronicleRec — Tencent compresses ultra-long behavior sequences in one pass into time-anchored Chronicle Tokens: recency-aware multi-granularity merging + causal encoder + multi-horizon masking + mask-and-predict pretraining, cacheable per user; seven-day online A/B confirms gains.
Single-Token — Indeed models candidate-job relevance as {1..5} ordinal classification, taking the expectation of the first token's probability as the score, and fine-tunes an SLM with a hybrid MSE+CE loss; online employer low-relevance -27.3%, keep rate +7.07%.
Query Suggestion — Alibaba uses intent-aware diversity rewards and query-level credit assignment for generative query suggestions, routing query-level quality signals to the corresponding tokens while sharing the slate-level diversity signal; online CTR, query quality, and intent coverage all improve.
LLM Agents and Retrieval-Augmented Generation
AURA — an end-to-end agentic system that reads production engagement logs to diagnose recommendation failure modes, then generates code-level improvements grounded in codebase context; preliminarily tested on production data from two consumer platforms at a major media streaming company, with an architecture transferable across platforms via a configuration layer.
MemRetriever — models long-term memory retrieval as agentic multi-step search, choosing parallel/serial search or reflection-and-denoising at each step with adaptive termination; ReAct-trace warm-start + GRPO; MemRetriever-4B-RL surpasses DeepSeek-v4-Flash on LongMemEval's main retrieval metric.
Pre-retrieval Query Clustering — Amazon estimates a per-query saturation point k* from the NDCG-k curve, then uses embedding clustering plus a mean-plus-variance rule to give cluster-level recommended retrieval depth, selecting top-k in constant time online; full-traffic F1 improves over 36%, low-complexity cluster token usage -14%.
  • Recommendation Systems
  • Weekly
  • Papers
  • AI Weekly 2026-W38AI Tech Daily - 2026-09-19
    Loading...