RecSys Weekly 2026-W39
2026-9-26
| 2026-9-26
字数 6007阅读时长≈ 16 分钟
type
Post
status
Published
date
Sep 26, 2026 05:31
slug
rec-weekly-en-2026-W39
summary
Industry papers came in denser than usual this week. TikTok, Kuaishou, Baidu, Google, LinkedIn, Meta, Spotify, and Walmart all shipped system papers with online A/B results, spanning the full stack from retrieval and pre-ranking to ranking and bidding. A second thread runs through evaluation and data quality — Netflix's counterfactual observability framework, Meta's synthetic data filtering, and an empirical audit that questions the evaluation protocol for LLM re-ranking. Thread one: continuous-space generative retrieval and unified cascades. X-Rec (ByteDance) abandons the discrete-token path of semantic IDs and instead learns the recommendation distribution directly in continuous item embedding space via flow matching. Inference throughput is 3.46× that of SID-AR, and vertical-content engagement on TikTok rose 4.1484%. OneTrans-V2 compresses retrieval, pre-ranking, and ranking into a single Transformer — GMV +9.74%, and 3.2× throughput on the same hardware. Both point the same way: the bottleneck in generative recommendation has shifted from "can we generate" to "how do we get both throughput and retrieval precision out of the generation path." Thread two: industrial systems correcting their own evaluation standards. Recall Ceiling finds that the oracle protocol commonly used in LLM re-ranking overestimates NDCG@10 by 92–95%, while real retrieval achieves only 2–19% Recall@100 — the ceiling locks the upper bound on re-ranking. FROST (Meta) attacks from the data side, using real-data gradients to anchor synthetic sample utility; filtering out 20–30% of synthetic data actually improves downstream performance. Work like this doesn't produce new models, but it changes how everyone reads everyone else's results. Thread three: MoE and parameter inheritance as engineering levers for scaling. IntBMoE (Alibaba AMap) decouples MoE participation, execution, and materialization through block-conditioned expert composition — serving hundreds of millions of users at 60ms latency
tags
Recommendation Systems
Weekly
Papers
category
Rec Tech Report
icon
📚
password
priority
1

Weekly Overview

Industry papers came in denser than usual this week. TikTok, Kuaishou, Baidu, Google, LinkedIn, Meta, Spotify, and Walmart all shipped system papers with online A/B results, spanning the full stack from retrieval and pre-ranking to ranking and bidding. A second thread runs through evaluation and data quality — Netflix's counterfactual observability framework, Meta's synthetic data filtering, and an empirical audit that questions the evaluation protocol for LLM re-ranking.
Thread one: continuous-space generative retrieval and unified cascades. X-Rec (ByteDance) abandons the discrete-token path of semantic IDs and instead learns the recommendation distribution directly in continuous item embedding space via flow matching. Inference throughput is 3.46× that of SID-AR, and vertical-content engagement on TikTok rose 4.1484%. OneTrans-V2 compresses retrieval, pre-ranking, and ranking into a single Transformer — GMV +9.74%, and 3.2× throughput on the same hardware. Both point the same way: the bottleneck in generative recommendation has shifted from "can we generate" to "how do we get both throughput and retrieval precision out of the generation path."
Thread two: industrial systems correcting their own evaluation standards. Recall Ceiling finds that the oracle protocol commonly used in LLM re-ranking overestimates NDCG@10 by 92–95%, while real retrieval achieves only 2–19% Recall@100 — the ceiling locks the upper bound on re-ranking. FROST (Meta) attacks from the data side, using real-data gradients to anchor synthetic sample utility; filtering out 20–30% of synthetic data actually improves downstream performance. Work like this doesn't produce new models, but it changes how everyone reads everyone else's results.
Thread three: MoE and parameter inheritance as engineering levers for scaling. IntBMoE (Alibaba AMap) decouples MoE participation, execution, and materialization through block-conditioned expert composition — serving hundreds of millions of users at 60ms latency, with UVCTR +2.4%. Inherit4Rec (Kuaishou) uses parameter inheritance for smooth dense-to-sparse transitions. Both answer the same question: how do you grow capacity without growing latency.

Generative Recommendation and Unified Ranking Architectures

X-Rec (ByteDance) makes a blunt argument: stop quantizing items into short token sequences. The SID-AR path has two problems — quantization error and the low throughput of autoregressive decoding. X-Rec instead learns the recommendation distribution directly over continuous item embeddings, uses flow matching to generate an embedding trigger, then hands off to ANN retrieval. Three designs support this: anchor conditioning splits generation into coarse semantic-region selection and fine-grained refinement; Riemannian flow matching aligns the generation trajectory with the hyperspherical geometry of item embeddings; and a late-interaction diffusion Transformer repeats the velocity-field estimate only at the final layer to contain compute. On TikTok's internal streaming benchmark, throughput is 3.46× SID-AR with matching retrieval quality. Two consecutive online launches delivered +4.1484% vertical-content engagement and +0.0111% overall engagement. That overall gain is tiny — the method's value today is concentrated in high-verticality scenarios. The lineage traces back to GPR's generative single-model framework — GPR used multi-level semantic ID mapping and heterogeneous hierarchical decoders end-to-end, while X-Rec simply removes the discrete mapping step and does coarse-to-fine generation in continuous space.
OneTrans-V2 takes the other path: don't change the generation paradigm, change the cascade structure. Traditional recommenders train and serve retrieval, pre-ranking, and ranking as three separate models — user behavior sequences get encoded three times, and optimization objectives stay isolated. OneTrans-V2 feeds all three stages into one Transformer, encodes the behavior sequence once as shared context, and retains stage-specific candidate features and computation. Three mechanisms deserve separate mention:
  • Sparse MoE scaling of the shared backbone: capacity grows with expert count, but activated computation stays bounded. A companion μP-style parameterization stabilizes training dynamics during upscaling.
  • DCGR (Decision-Conditioned Generative Retrieval): merges multiple retrieval channels — previously split by business objective — into one. It first predicts a decision prefix describing the imminent interaction, then generates items conditioned on it. Business objectives become conditions on the generation process rather than separate models.
  • SNT (Sequence-Native Training): organizes training around each user's lifetime behavior sequence, amortizing sequence encoding cost across exposures.
Deployed online across all three stages: GMV +9.74%, and 3.2× throughput on the same hardware with a co-designed serving stack. That 3.2× matters more than X-Rec's 3.46× — it comes from architectural consolidation itself, not from optimizing within a single stage.
UNIQUE (Baidu) also unifies retrieval and ranking, but with a more specific entry point: hierarchical quantization instability, and the information loss from separating retrieval and ranking. It replaces hierarchical quantization with single-layer flat quantization, fuses generative code-based retrieval and target-aware ranking into one early-fusion architecture trained end-to-end, and adds a balanced quantization mechanism to ease codebook imbalance and improve long-tail representations. Deployed across three scenarios on Mobile Baidu — home feed, discovery, and short video — total watch time rose 0.96% and total distribution volume 1.08%, with P99 latency at 89ms and online inference MFU at 44.23%. New users and highly active users improved more — consistent with the motivation that flat quantization helps long-tail representations.
CMRec (Alibaba) tackles cross-country recommendation. E-commerce platforms typically maintain disjoint user and item ID spaces across markets, so the shared anchors that traditional cross-domain methods rely on simply don't exist. Generative recommendation eases part of this through a shared token space, but existing methods still split behavior sequences strictly by country — knowledge transfer happens only at the parameter level. CMRec borrows the code-switching idea from multilingual NLP and pushes transfer down to the data level: first learn a shared semantic codebook from multimodal content and cross-country behavioral co-occurrence, then use it to synthesize mixed-country sequences. Token-level substitution must satisfy both static constraints (content) and dynamic constraints (price, audience, popularity). Finally, a context-aware loss reweights each sample by its plausibility within the current sequence. On real multi-country e-commerce data plus online A/B: ad revenue +1.77%, orders +2.64%, and gains in data-sparse countries don't come at the expense of data-rich ones. The precursor here is OpenOneRec, which used on-policy distillation plus Rec-RL in a two-stage alignment to push 1.7B/8B models to an average +26.8% Recall@10 across 10 Amazon datasets. CMRec cares not about single-market performance but about injecting data-level supervision across markets.
Retrieval-Grounded Credit Assignment (Microsoft) targets the credit-assignment gap in reasoning-enhanced generative recommendation. These models first generate a text reasoning trace, then beam-search decode SIDs, trained with group-relative policy optimization and rewarded on exact SID match. With a large catalog, that reward is extremely sparse, producing two failure modes: when every rollout in a group misses, advantage is identically zero and there's no learning signal at all; and different traces get identical advantage as long as the final SID matches, even when reasoning quality differs wildly. The paper's approach structures the trace into three segments — history summary, interest hypotheses, and final SID — and uses a frozen retriever to execute each hypothesis as a catalog query. Each hypothesis becomes independently verifiable rather than judged indirectly through the final SID. Any query hitting the target within top-K rewards that rollout, and per-query hit indicators localize reward to specific hypotheses. Credit therefore lands at the span level: only individually hitting hypotheses receive positive advantage, and the retrieval channel never updates the final SID span. Consistent gains across three Amazon Reviews datasets. Oracle analysis on Video Games further shows that selecting target-relevant queries from generated interests improves both recall and ranking — meaning interest quality itself still has unrealized headroom. Related work: DualGR routes long- and short-term interests through dual branches with search-based SID decoding, while Gryphon jointly trains item-level scoring components to resolve the mismatch between SID likelihood and relevance objectives. All three sit on the same axis: how the SID generation process aligns with true relevance.

Ad Bidding and Agent Optimization

OneBid (Kuaishou) aims to build a foundation model for auto-bidding. Existing methods evolved from rule-based controllers to reinforcement learning to Decision Transformers, but none match today's oCPX paradigm — registration, purchase, and other heterogeneous scenarios each get their own model, the pipeline is fragmented, and cross-scenario modeling is entirely unexplored. Unifying into a single model poses three difficulties: multi-objective control, capacity scaling under strict latency, and safe offline policy improvement. OneBid's answers:
  • Dual atomic signal conditioning: expands DT's single Return-to-Go into two — Return-to-Go for conversion value, Cost-to-Go for cost ratio. A value-aware regularizer constrains next-action prediction.
  • Sequence-level MoE: shared experts encode cross-scenario knowledge, sparse routed experts capture scenario-specific patterns, ensuring capacity grows consistently with model size and data volume at low latency.
  • CROP (Critic-guided Relative Offline Policy): learns a critic to score candidate actions relatively within groups, avoiding the online exploration risk of GRPO-style fine-tuning while constraining policy shift to reduce OOD risk.
Fully deployed online: oCPX Ads overall ADVV +2.2%, with a peak of +13.1% in ROAS scenarios. That gap is itself informative — the unified model gains most where objectives are clear and feedback dense, and gains get diluted where objectives are diffuse.
ADAPT (Alibaba Taotian) also targets auto-bidding personalization, but focuses on advertiser profiles. Profiling has long been validated in recommendation, yet porting it to bidding is hard for three reasons: clean profiles are hard to extract, public and private information are hard to model jointly, and profile updates plus cold-start adaptation are painful. ADAPT uses two-stage training: stage one extracts clean static and dynamic profiles via contrastive learning over an advertiser memory bank; stage two decouples dynamic profiles into public and private parts, then conditions the bidding policy on these plus the static profile. After training, neither profile construction for new advertisers nor profile updates for existing ones require retraining. So far validated only on a large-scale offline benchmark — no online numbers. Its focus differs from PRO-Bid, which handles constraint satisfaction via Constraint-Decoupled Pareto representations and counterfactual regret optimization; ADAPT handles "who should the bidding policy be conditioned on."
A Pinch of SFT, A Dash of RL (Amazon) asks how to allocate SFT versus RL in long-horizon advertising agents. Enterprise analytics agents perform multi-step tool calls over distributed business data — retrieval, reasoning, API calls, code execution — and must adapt to intermediate observations. SFT calibrates tool syntax and teacher-supported behaviors; RL explores reward-reachable behaviors beyond demonstrations. But applying RL uniformly disrupts already-calibrated skills. The paper's observation: checkpoint trajectories can be partitioned post hoc into three regimes — Imitation (SFT already captures reliable teacher behavior), Lift (both stages help), and Discovery (useful reward-observable behavior lies outside reliable teacher support). More valuable: this partition can be used prospectively. Based on teacher support and reward-observable headroom, route features to SFT-only, SFT-then-RL, additional RL, or environment fixes first. In follow-up experiments across 18 features, this diagnostic predicted the correct trajectory direction in 15 cases. On GPT-OSS 120B, targeted SFT followed by RL produced positive point estimates on 7 of 8 advertising skills, with paired 95% confidence intervals excluding zero on 5 — the largest gain on non-disclosure at +11.27 points (95% CI [+9.72, +12.82]). SME audit also found targeted RL cut standard information leakage from 11.8% to 2.9% and adversarial leakage from 22.9% to 6.8%, while largely preserving actionability (86.2% → 85.7%). Under shared reward and optimization configuration, targeted RL raised the seven-skill mean from +1.62 to +3.57 while using 43% less incremental RL compute.
FROST (Meta) addresses synthetic data selection. Existing methods emphasize fidelity or diversity rather than what the learner needs at its current stage. FROST uses gradient feedback from real training data to estimate synthetic sample utility, with two core decisions: when to filter — calibrate against batch utility and recent history, triggering filtering only on out-of-band batches; and what to keep — sample selection also happens only within out-of-band batches. The whole process needs no external verifier and no held-out validation set. On two public benchmarks — image classification and text-to-SQL fine-tuning — filtering out 20–30% of synthetic data improved real-task performance. Applied to Meta's large-scale ad re-ranking system, it delivered substantial gains over a highly optimized production baseline. Specific A/B numbers aren't in the abstract, but "filtering data out makes things better" is a direct methodological reminder for teams relying on synthetic data augmentation.

LLM-Enhanced Re-Ranking and User Modeling

Start with a paper whose conclusions may not be popular. Recall Ceiling audits the evaluation protocol for LLM recommendation re-ranking. Many LLM re-ranking works adopt an oracle protocol — injecting ground-truth items into the candidate list, or scoring them against sampled negatives. Across three major Amazon datasets, the paper shows this protocol overestimates true NDCG@10 by 92–95%. The cause is a deterministic upper bound: across eight datasets and three domains, real retrieval covers only 2–19% of relevant items at K=100. Under leave-one-out evaluation, $\mathbb{E}[\mathrm{NDCG}@k] \leq \mathrm{Recall}@|W_\pi|$, where $W_\pi$ is the re-ranking candidate window. Under real retrieval conditions, none of the optimization strategies tested — prompt engineering, model scaling across a 168× parameter range, sequential models, supervised neural re-rankers, LoRA fine-tuning, hybrid retrieval, score-aware prompting, LLM+CF fusion — significantly beat a collaborative filtering baseline. Text-aware retrieval improved recall on one dataset but left end-to-end NDCG unchanged; feeding upstream CF scores to the LLM mainly made it restate the CF ranking. The paper therefore proposes RAEP (Recall-Aware Evaluation Protocol): first determine which recall regime retrieval is in, then decide whether re-ranking is worth evaluating. In the low-recall regime the paper measured, improving retrieval beats increasing re-ranker complexity — though the paper itself notes this ordering may not hold in production systems with higher recall, richer features, or online feedback. Related work Adaptive Re-Ranking uses a utility function for query routing, cutting median latency 1.15–53× across multiple IR datasets with nDCG@10 changes between -17.5% and +4.0% — essentially also answering "when is re-ranking not worth doing."
LLM User Profiling (Comcast) asks a very practical question: LLM-generated user profiles cost far more than aggregation methods, so when is that money worth spending. The paper factorially crosses four semantic user profiling strategies — representation type (aggregate vs LLM-generated) × temporal handling (whether recent and historical behavior are decoupled) — evaluated on a real production streaming dataset. Comparisons span user behavior types, accuracy and non-accuracy dimensions, and controlled temporal window settings. At the conclusion level, this isn't a new method but evidence for engineering decisions — practitioners need systematic comparisons like this more than yet another profiling model. Related work HyMiRec takes a hybrid route: a lightweight recommender extracts coarse-grained interest embeddings, an LLM recommender captures fine-grained ones, and a cosine-similarity residual codebook compresses user history. Sci-Surf combines LLM profiling with multimodal summarization, with online evaluation showing +10.4% prediction alignment.
COPE (Alibaba Qwen) handles continual personalization under sparse feedback. Aligning LLMs to canonical values homogenizes responses, failing to cover diverse user preferences; training-free methods burn precious context window on prompt engineering, and training-based methods go static once trained. COPE assigns each user a learnable personalized embedding and integrates preference capture, self-evaluation calibration, and personalized response optimization into a single update step. The key design uses self-evaluation to generate a proxy reward — the model keeps updating even when explicit user feedback is absent. It consistently outperforms training-free and training-based baselines under sparse feedback, and complements Retrieval-Augmented Prompting. This sits in the same lineage as Aligning LLMs for Controllable Recommendations's two-stage SL+RL alignment — both replace explicit labels with learning signals. The difference: COPE sources its signal from model self-evaluation, removing dependence on labeled feedback entirely.
Robust Fusion (Spotify) examines the side effects of stuffing behavioral statistics into LLM prompts. QSS (Query Slice Stats) is an interaction-derived behavioral feature summarizing historical success for query-candidate pairs. Injecting QSS directly improves ranking when the feature is available, but robustness drops when it's removed — classic shortcut learning, where the model learns to rely on historical signals and ignore semantic and user-context patterns. The fix is deterministic dual-sample feature-dropout training: each sample appears twice, once with QSS and once without. Offline, QSS injection improves ranking quality by 13.3% when available, and dual-sample training preserves that gain while improving 4.0% over naive QSS training under QSS-removal evaluation. In online tests, both QSS-aware variants raised search success rate by roughly 2% — note that aggregate testing can't distinguish dual-sample from feature-only training, and the cold-start comparison only aligns directionally with offline results. This honest negative result is worth recording.
BoundaryMORPH (AWS AI Labs + Purdue) targets a structural mismatch in RAG. Open-ended queries are increasingly "diffuse," requiring large document sets to fit into a limited LLM context window. Systems use fast dual-tower retrieval, then score with a more expensive cross-encoder, but the CE budget B is tightly latency-constrained and typically smaller than context capacity k. Standard re-ranking therefore structurally wastes compute — verifying obvious top candidates while ignoring relevant documents ranked lower initially. BoundaryMORPH uses a Gaussian process to treat dual-tower initial ranking as a structural prior, concentrating CE calls on resolving boundary membership of the top-k set rather than finding a single most-relevant document. Information from each CE score propagates to unscored documents, maximizing budget utility. It achieves SOTA across multiple models and datasets, with nCG@100 5.4 above the strongest baseline. Precursors: GRank uses a target-aware generator plus lightweight ranker, improving Recall@500 by over 30% on public benchmarks and production corpora with P99 QPS 1.7× that of tree/graph retrievers. BAGEL combines LLM relevance scoring with Gaussian-process Bayesian active learning, beating LLM re-ranking methods at equal LLM budget. All three share a judgment: in budget-constrained re-ranking, deciding *where to spend queries* matters more than deciding *which model to use*.
Two evaluation-and-infrastructure papers also belong here. Counterfactual Observability (Netflix) models recommender observability as a counterfactual measurement problem — estimating what the recommender would do, and what interactions would follow, absent a given piece of content or model decision. The paper distills three observability principles for content creators and model developers, plus measurement methodology covering coverage-bias reduction, relativity, and incrementality, applicable to both single-stage and cascaded recommenders, and already deployed across multiple Netflix production systems. Raw interaction signals (plays, clicks) conflate content quality, model behavior, presentation bias, and audience reach — this framework addresses the attribution problem.
RAILS (Zendesk) makes LLM clustering production-ready. Using an LLM as clusterer is hard at production scale — prompts can't hold the entire label space, and serial per-document processing lacks throughput. RAILS turns clustering into a loop over a growing label pool, scaling through document batching and bounded concurrency. Across six public benchmarks it beats the strongest LLM clustering methods on average: accuracy 51.2% → 59.3%, NMI 67.2% → 74.8%, ARI 45.4% → 54.7%. It replaced the traditional HDBSCAN stage in a SaaS ticket topic-discovery pipeline, delivering higher clustering quality, prompt-driven transparent control, and stateful incremental runs. Related work SHERLOCK achieved 82% expert acceptance and a 386.7% increase in daily survey throughput in JD.com's online A/B, following the same two-stage strategy of domain knowledge base plus retrieval augmentation.

Engineering Scaling of Large-Scale Retrieval and Ranking Systems

IntBMoE (Alibaba AMap) offers the clearest analytical framework for MoE this week. For a single token, three quantities previously couldn't be set independently: participation (how many experts contribute knowledge to the output), execution (how many actually compute, determining compute cost), and materialization (how many expert parameter sets must be built and stored, determining memory cost). Sparse routing suppresses execution and materialization but also participation; dense output-mixing restores full participation but execution grows with expert count; parameter-merging keeps execution at one expert but materialization grows with the number of routing decisions. IntBMoE decouples all three via block-conditioned expert composition: blocks come from a small learned codebook, one entry each; a lightweight hypernetwork per layer merges all expert bases in that layer's pool into a single composed expert. Participation is full because the whole pool participates; execution is sparse because the router sends tokens to only a few blocks; materialization is bounded because block count is determined by the codebook, not the input. Dual-Path Residual Gating then couples two independently composed paths via multiplicative gating. It consistently beats sparse and dense MoE baselines on image classification, with generalization validated on language modeling and sequential recommendation. Deployed in AMap's generative recommendation system, serving hundreds of millions of users within a 60ms latency budget, with online A/B showing +2.4% relative UVCTR.
Inherit4Rec (Kuaishou) addresses the cost of model scaling. Repeatedly training larger dense models from scratch demands enormous data and time, and the added compute conflicts with strict industrial serving budgets. Parameter inheritance is a viable path, but existing methods target static corpora and lose accuracy on dynamically evolving recommendation data. Inherit4Rec supports two paths: D2D (Dense-to-Dense) uses hybrid growth plus asymmetric training, preserving the forward function and update continuity during expansion; D2S (Dense-to-Sparse) builds an SMoE via co-activation-aware partitioning and load-balancing loss, retaining dense-model capability while promoting balanced expert activation. On KuaiRand-1K and an industrial short-video dataset, both transformations beat the compared inheritance baselines on all prediction objectives. Related work: LoopCTR takes another compute-saving path — recursively reusing shared model layers to add training-time compute, decoupling compute growth from parameter growth, training a multi-loop-inference zero-loop policy. Disaggregated Multi-Tower approaches from data-center topology, using semantics-preserving tower transformations plus hierarchical feature interaction to achieve up to 1.9× speedup on data centers without accuracy loss.
Lightweight Ranking Heads (Google/YouTube) solves not model quality but experiment velocity. Adding a new prediction task to a production multi-task ranking model often becomes the bottleneck — it may conflict with existing tasks, and it requires retraining the backbone and downstream models plus tuning the reward combination formula. Light Heads allows dynamic injection of new tasks into an existing multi-task ranking model, strictly isolating new tasks via stop-gradient and stateless daily training, and using centralized configuration so the same set of heads attaches to multiple models simultaneously — unlocking faster training-data generation and downstream co-training. After YouTube-scale deployment, multi-task experiment iteration cycles dropped from weeks to days. The engineering value is clear: "add a task" gets downgraded from a model-engineering problem to a configuration problem.
ScalarLens (Ant Group) makes a direct critique of numerical feature embedding: a scalar has only one representation, and that premise conflates "where the value falls" with "what it means for the current sample." On the Criteo validation set, the paper observes that even after removing additive main effects, the same numerical interval carries residual click evidence with opposite signs in categorical versus numerical contexts. Production pipelines amplify the mismatch — externally normalized features require transformations and statistics to stay synchronized between training and serving. ScalarLens's position is to redefine numerical embedding as a measurement problem: coordinates belong to the value itself, using a monotone local grid to construct stable coordinates from the focal scalar; predictive response belongs to the value in context, using bounded low-rank dynamics to generate contextual response without moving that coordinate or replacing categorical tokens and the CTR backbone. The main evaluation spans 1,539 runs across 19 representations, 3 datasets, 9 backbones, and 3 seeds; ScalarLens ranks first in 25 of 27 settings and second in the other two. Matched ablations show scale correction, extra local capacity, and generalized conditioning all fail to reproduce the gain; a full rerun under shared normalization still substantially outperforms DEER, DAES, and NaryDis. For lineage, see DEI²N on modeling immediate-interest dynamics in trigger-based recommendation, and INFNet using hub tokens and aggregate-and-broadcast for linear-complexity feature interaction — the latter deployed in a commercial online ad system with revenue +1.587% and CTR +1.155%.
CC Retriever (LinkedIn) handles graph features in pre-ranking. On LinkedIn Feed, content from a user's connections (friends and follows) accounts for over 70% of impressions and engagement, and interaction signals in the professional knowledge graph span first- and second-degree connections — including "stranger viral" content, where a first-degree connection reacts to, comments on, or reposts content they didn't author. This fan-out pushes the candidate index past a billion, filtered to roughly tens of thousands that must be scored within a 120ms P99 latency budget. CC Retriever scores these candidates with a full-depth ranking model on GPU, centered on a sorted-search GPU primitive that joins dense graph affinity features (viewer to author) with document-level features in 5–10ms at runtime. After moving to GPU-served scoring, ranking model parameters scaled 50×, and online content consumption time rose 2.5% — a notably high figure for LinkedIn Feed experiments. For reference: FSCD uses learnable feature selection for the effectiveness-efficiency tradeoff in pre-ranking, and ASH/ASMOL proposes all-scenario metrics — pre-ranking evaluation standards are themselves an open problem.
Hybrid GPU-CPU Retrieval (Meta) characterizes what it calls the "personalization-scale paradox": fitting the full serving inventory into GPU memory is too expensive, while CPU compute can't run the same heavy interaction models. The solution is orchestration, not a new model class. The GPU path does deep fusion — retrieval plus interaction pre-ranking — over a curated online pool of roughly a billion documents. The CPU path does lightweight personalized scoring over an independent online inventory roughly 20× larger. Each request can use one path or both, with candidates deduplicated and sharing downstream ranking. Full-system A/B against a legacy CPU-only configuration improved both model-scoring relevance metrics and substantive engagement; per-path experiments showed positive gains within each deployment scope. Retrieval logs show the two paths contribute structurally complementary candidates. Related work: Semantic Search at LinkedIn uses an LLM relevance judge plus multi-teacher distillation into a small model, improving ranking throughput over 75× at fixed latency. BEBR (Tencent) starts from binary embeddings, using cyclic binarization and SDC distance computation to save 30%–50% index cost with near-zero loss.
MuSeR (Baidu) does long-sequence retrieval on the already-deployed MGS system. Industrial systems typically truncate history to a few hundred actions under latency and memory constraints, underusing long-term interests; users also pursue heterogeneous intents across news, Q&A, short video, and other modalities, which sparse ID embeddings struggle to represent. MuSeR has three components: hierarchical temporal compression — recent actions stay at full resolution while earlier segments are progressively pooled, fitting single-user histories of $10^4$–$10^5$ scale into a fixed serving budget; decoupled multi-query interest extraction with orthogonal regularization; and multimodal semantic alignment — augmenting sparse item IDs with LLM-distilled text summaries. The deployment side pairs asynchronous user representation refresh, adaptive caching, and hierarchical beam-search retrieval across heterogeneous hardware. Online A/B across three Baidu APP scenarios — home feed, discovery, and short video: DAU +0.26%, total time +0.89% (both p<0.05), with serving latency and cost both down. The paper itself says this isn't a new modeling primitive but a system-level integration that makes long-term, multi-interest, multimodal modeling jointly deployable — an accurate framing. Related work OrDA uses orthogonal regularization to constrain interest and habit latent spaces to be geometrically perpendicular and performs causal intervention at inference, with online UCTR +5.64%.
EvoPilot (Meta) belongs to the emerging class of work on running LLM agents for long-horizon online autoresearch. Autoresearch loops on self-contained programs iterate in minutes, but online autoresearch spans asynchronous systems, hour-scale variants, and week-scale campaigns, directly affecting product. A completed experiment can still support a wrong conclusion — code changes are no-ops, data windows leak, evaluator semantics drift, two arms traverse different serving funnels. EvoPilot is a human-gated method: role-specialized agents execute each round through versioned domain skills and typed adapters, persistent records store experiments and failures, and deterministic checks enforce recorded lessons. The paper studies a 37-day campaign against Meta Video Deep Dive's retrieval system, covering seven directions over a hundreds-of-millions video index refreshed hourly. Prior manual experiments failed to establish gains from an interaction head; a primitive autoresearch attempt revisited the direction but misattributed a 22-percentage-point offline hit-rate drop to that head. With EvoPilot, human-gated verification traced the drop to a pre-existing evaluation defect producing output depths of 3,000 and 600. After the fix, matched comparison measured a 3.20-percentage-point offline gain. Post-hoc replay and mutation testing rejected invalid comparisons and accepted valid ones. Persistent state recovered an interrupted round, and artifact reuse saved roughly 5 GPU hours. A separate seven-day randomized online evaluation estimated a 0.66% relative GSRR improvement on the VDD slice. Related directions: Applying EBR to Airbnb Search on multi-stage user journey modeling, and DualAgent-Rec using an LLM as coordinator in dual-agent constrained optimization — the latter achieving 100% constraint satisfaction and 4–6% Pareto hypervolume improvement on Amazon Reviews 2023.
GradCIR (Walmart) handles a special query class in e-commerce visual search — composed image retrieval, where users upload an image plus a text description of the desired modification (change color, change style). Existing CIR methods treat relevance as binary and train on triplets with a single positive, but in real catalogs many candidates only partially satisfy the query, and ranking along that partial-match spectrum determines the experience. GradCIR's methodology has three parts: VLM-generated queries (object detection plus modifier synthesis) and 4-level relevance labels, requiring no human annotation; an iterative relevance-feedback loop that mines hard negatives from the in-training retriever to expand the training set; and a hierarchy-aware angular objective that trains the retriever directly on graded labels rather than collapsing them to binary. Trained on 3.5M graded pairs (from Walmart's raw catalog) on a PaliGemma2 dual tower. A controlled graded-vs-binary ablation isolates the effect of supervision granularity, with NDCG@10 improving 4.9%–5.9%; the same recipe applied to other multimodal encoders yields up to 8.5% NDCG@10 improvement on early-fusion backbones. On FashionIQ, average recall is 0.6703, slightly ahead of the strongest peer-reviewed supervised baseline and matching or exceeding all published CLIP-L-class zero-shot CIR methods. The system already serves online visual search traffic in Walmart production; the abstract gives no online A/B numbers.

Directions to Watch

Evaluation protocols are becoming research objects themselves. Three papers this week question existing evaluation standards from different angles. Recall Ceiling demonstrates that the oracle protocol overestimates NDCG@10 by 92–95%, and that real retrieval's 2–19% recall forms a deterministic upper bound. EvoPilot documents a complete misattribution of a 22-percentage-point drop and how deterministic checks intercepted it. Robust Fusion honestly reports that online aggregate testing can't distinguish dual-sample from feature-only training. All three point to the same thing: as model complexity rises, bias from evaluation defects can exceed the magnitude of methodological improvement itself. For teams doing LLM re-ranking, RAEP's "determine the recall regime before evaluating" workflow is worth borrowing directly.
The route dispute between continuous-space generation and discrete SIDs is unresolved. X-Rec uses flow matching to sidestep quantization error and autoregressive low throughput, gaining 3.46× throughput. OneTrans-V2's DCGR still does decision-conditioned generation within the SID framework. Retrieval-Grounded Credit Assignment focuses on repairing the credit-assignment gap in RL training under the SID route. X-Rec's overall gain is only +0.0111%, indicating continuous space currently wins only in high-verticality, semantically concentrated scenarios. Worth tracking: whether work like Gryphon — jointly training item-level scoring to resolve the mismatch between SID likelihood and relevance objectives — changes this landscape.
MoE capacity scaling is shifting from "stack more experts" to "decouple cost terms." IntBMoE separates and controls participation, execution, and materialization, serving hundreds of millions of users within a 60ms budget on AMap. Inherit4Rec uses parameter inheritance for smooth dense-to-sparse transitions. OneBid's sequence-level MoE serves cross-scenario bidding. All three answer the same constraint — capacity must grow, but latency and memory can't grow proportionally. Next question worth watching: whether this decoupling framework generalizes to the sequence-length dimension, not just the expert dimension.

Paper Roundup

Generative Recommendation and Unified Ranking Architectures
X-Rec — ByteDance learns the recommendation distribution directly in continuous item embedding space via flow matching and generates embedding triggers for ANN retrieval; inference throughput 3.46× SID-AR, TikTok vertical engagement +4.1484%, overall +0.0111%.
OneTrans-V2 — Unifies retrieval/pre-ranking/ranking into a single Transformer cascade, introduces DCGR to unify multi-objective retrieval channels and SNT to amortize lifetime sequence encoding, with sparse MoE + μP scaling; GMV +9.74%, 3.2× throughput on the same hardware.
UNIQUE — Baidu unifies generative code-based retrieval and target-aware ranking via single-layer flat quantization plus early fusion; across three Mobile Baidu scenarios, total watch time +0.96%, total distribution volume +1.08%, P99 89ms.
CMRec — Alibaba transfers code-switching from multilingual NLP to cross-country generative recommendation, using a shared semantic codebook to synthesize dual-constrained mixed-country sequences; ad revenue +1.77%, orders +2.64%.
Retrieval-Grounded Credit Assignment — Microsoft splits the reasoning trace into history summary, interest hypotheses, and final SID, executing each hypothesis with a frozen retriever for span-level credit assignment; consistent gains across three Amazon Reviews datasets.
Ad Bidding and Agent Optimization
OneBid — Kuaishou's first unified oCPX auto-bidding foundation model, expanding DT's single Return-to-Go into Return-to-Go + Cost-to-Go dual signals, with sequence-level MoE and CROP offline policy optimization; oCPX Ads overall ADVV +2.2%, peak +13.1% in ROAS scenarios.
ADAPT — Alibaba Taotian uses contrastive learning to extract clean static/dynamic advertiser profiles, decoupling dynamic profiles into public and private parts, supporting training-free profile construction for new advertisers; consistently ahead on a large-scale auto-bidding benchmark.
A Pinch of SFT, A Dash of RL — Amazon proposes an Imitation/Lift/Discovery three-regime diagnostic framework to prospectively route SFT/RL allocation; on GPT-OSS 120B, 7/8 advertising skills positive, non-disclosure +11.27 points, targeted RL saves 43% incremental compute.
FROST — Meta uses real-data gradients to anchor synthetic sample utility, deciding when to filter and what to keep only on out-of-band batches; filtering 20–30% of synthetic data improves real-task performance, already used in large-scale ad re-ranking.
LLM-Enhanced Re-Ranking and User Modeling
Recall Ceiling — Reveals that the oracle protocol for LLM re-ranking overestimates NDCG@10 by 92–95%, with real retrieval Recall@100 at only 2–19%; proposes the RAEP evaluation protocol; in low-recall regimes, improving retrieval beats upgrading the re-ranker.
BoundaryMORPH — AWS AI Labs + Purdue use a Gaussian process to treat dual-tower ranking as a structural prior, concentrating cross-encoder budget on top-k boundary membership decisions; nCG@100 5.4 above the strongest baseline.
Robust Fusion — Spotify uses deterministic dual-sample feature-dropout training to mitigate shortcut learning on QSS behavioral features; offline ranking +13.3% when QSS available, +4.0% under QSS removal, online search success rate ~+2%.
COPE — Alibaba Qwen assigns each user a learnable personalized embedding and uses self-evaluation to generate a proxy reward supporting continual updates under sparse feedback; consistently beats training-free and training-based baselines.
LLM User Profiling — Comcast factorially evaluates four semantic user profiling strategies (aggregate vs LLM-generated × temporal decoupling or not), answering when LLM profiling justifies its extra cost.
Counterfactual Observability — Netflix models recommender observability as a counterfactual measurement problem, proposing three principles for creators and model developers plus measurement methodology covering coverage-bias reduction, relativity, and incrementality.
RAILS — Zendesk turns LLM clustering into a retrieval-augmented loop over a growing label pool with batching and bounded concurrency; across six benchmarks ACC 51.2% → 59.3%, already replacing the HDBSCAN stage in a production pipeline.
Engineering Scaling of Large-Scale Retrieval and Ranking Systems
IntBMoE — Alibaba AMap uses block-conditioned expert composition to decouple MoE participation, execution, and materialization; serving hundreds of millions of users at 60ms latency, online relative UVCTR +2.4%.
Inherit4Rec — Kuaishou proposes a parameter inheritance framework: D2D uses hybrid growth + asymmetric training to preserve the forward function, D2S uses co-activation-aware partitioning + load-balancing loss to build an SMoE; comprehensively ahead on KuaiRand-1K and an industrial short-video dataset.
Lightweight Ranking Heads — Google/YouTube uses stop-gradient and stateless daily training to support dynamic injection of new tasks into existing multi-task ranking models; multi-task experiment iteration cycles drop from weeks to days.
ScalarLens — Ant Group redefines numerical embedding as a measurement problem: monotone local grids construct stable coordinates, bounded low-rank dynamics generate contextual response; ranks first in 25 of 27 settings across 1,539 runs.
CC Retriever — LinkedIn uses a sorted-search GPU primitive to join dense graph edge features with document features in 5–10ms, enabling full-depth ranking models on GPU for pre-ranking; parameters scaled 50×, content consumption time +2.5%.
Hybrid GPU-CPU Retrieval — Meta solves the personalization-scale paradox with a GPU high-depth fusion path (billion-scale pool) plus a CPU high-breadth lightweight path (~20× inventory); full-system A/B improves relevance and substantive engagement over a legacy CPU-only configuration.
MuSeR — Baidu does long-sequence retrieval on the deployed MGS system using hierarchical temporal compression, multi-query interest decoupling, and LLM text-summary multimodal alignment; across three Baidu APP scenarios, DAU +0.26%, total time +0.89%.
EvoPilot — Meta proposes a human-gated long-horizon online autoresearch method; in a 37-day campaign it traced a 22pp offline drop to an evaluation defect, with +3.20pp offline after the fix; a seven-day randomized online experiment showed +0.66% relative GSRR.
GradCIR — Walmart shifts composed image retrieval from binary relevance to 4-level graded supervision, using VLM auto-labeling + iterative hard-negative mining + a hierarchy-aware angular objective; NDCG@10 improves 4.9%–5.9%, already serving online visual search traffic.
Other
RLVR² — Baidu's ERNIE team converts multi-dimensional rubric scores into within-group ordinal outcomes, recovering latent utility from comparison matrices and fusing into a single training signal, avoiding heterogeneous rubric scale calibration; consistently beats rubric-based baselines across 3 model scales and 16 benchmarks.
  • Recommendation Systems
  • Weekly
  • Papers
  • AI Weekly 2026-W39AI Tech Daily - 2026-09-26
    Loading...