RecSys Weekly 2026-W40
2026-10-3
| 2026-10-3
字数 4051阅读时长≈ 11 分钟
type
Post
status
Published
date
Oct 3, 2026 05:31
slug
rec-weekly-en-2026-W40
summary
Three themes run through recommender systems research this week. All three share a common shift: from "the method works" to "the method holds up in production." Theme 1: Generative recommendation is moving from point innovations to end-to-end engineering. ByteDance's GEAR for Douyin Ads puts the tokenizer, generator, and reranker into a single differentiable framework, tackling the coupled bottlenecks of representation collapse and item collision head-on. Tencent's KuaFu compresses billion-scale user behavior into item-level tokens. GRP v0.1 uses a single encoder-decoder to handle retrieval, ranking, and reward modeling at once. All three point at the same problem: the bottleneck in generative retrieval is no longer model capacity — it's tokenizer stability and serving cost. Theme 2: The payoff boundary for LLM-based user understanding is coming into focus. A controlled experiment on a production streaming platform finds that LLM-generated profiles beat aggregate profiles only for exploratory users. On habitual consumers — roughly four-fifths of the base — they're actually weaker. This is counterintuitive but actionable. It directly determines that profile strategy should be selected dynamically by consumption mode, not swapped in globally. Theme 3: Training infrastructure and retrieval architecture are being pulled apart as standalone topics. Meta's ETT turns training lifecycle overhead into attributable metrics, lifting fleet-level ETT% from around 80% to above 90%. TikTok's HELIX and LinkedIn's ESP make structural changes to ranking and retrieval — asymmetric scaling and multi-objective decoupling. Each of these looks incremental on its own. Together they show that the main source of gains in industrial systems is shifting from "swap the model" to "change the structure."
tags
Recommendation Systems
Weekly
Papers
category
Rec Tech Report
icon
📚
password
priority
1

Weekly Overview

Three themes run through recommender systems research this week. All three share a common shift: from "the method works" to "the method holds up in production."
Theme 1: Generative recommendation is moving from point innovations to end-to-end engineering. ByteDance's GEAR for Douyin Ads puts the tokenizer, generator, and reranker into a single differentiable framework, tackling the coupled bottlenecks of representation collapse and item collision head-on. Tencent's KuaFu compresses billion-scale user behavior into item-level tokens. GRP v0.1 uses a single encoder-decoder to handle retrieval, ranking, and reward modeling at once. All three point at the same problem: the bottleneck in generative retrieval is no longer model capacity — it's tokenizer stability and serving cost.
Theme 2: The payoff boundary for LLM-based user understanding is coming into focus. A controlled experiment on a production streaming platform finds that LLM-generated profiles beat aggregate profiles only for exploratory users. On habitual consumers — roughly four-fifths of the base — they're actually weaker. This is counterintuitive but actionable. It directly determines that profile strategy should be selected dynamically by consumption mode, not swapped in globally.
Theme 3: Training infrastructure and retrieval architecture are being pulled apart as standalone topics. Meta's ETT turns training lifecycle overhead into attributable metrics, lifting fleet-level ETT% from around 80% to above 90%. TikTok's HELIX and LinkedIn's ESP make structural changes to ranking and retrieval — asymmetric scaling and multi-objective decoupling. Each of these looks incremental on its own. Together they show that the main source of gains in industrial systems is shifting from "swap the model" to "change the structure."

Unified Architectures and Engineering Deployment for Generative Recommendation

The core tension in putting generative retrieval into production is holding three things stable inside one differentiable framework: tokenizer stability, candidate distinguishability, and affordable serving. This week's three deployment papers each address a different piece.
GEAR (ByteDance) — an end-to-end generative retrieval framework for Douyin Ads. It handles two bottlenecks that the literature usually treats separately: representation collapse, where the item tokenizer converges to degenerate results under continuous distribution drift, and item collision, where different items in a large candidate pool share the same token sequence. The two are coupled — expanding codebook capacity to ease collisions simultaneously worsens collapse.
The technical solution has two layers. For collapse, GEAR proposes BasisVQ: reparameterizing the codebook with orthogonal bases so gradients are shared globally, plus a rigid rotation on the latent space. The authors report that this needs no ad-hoc heuristics — the gradient dynamics stabilize on their own. An extended prefix-aware BasisRQ improves codebook expressiveness at the same asymptotic time complexity, at acceptable cost. For collisions, GEAR attaches a context-conditioned reranking head to the generation process, disambiguating colliding items with minimal extra compute.
This design contrasts with other approaches in the field. GPR took the route of a unified input schema plus a heterogeneous hierarchical decoder in a single end-to-end model. MDGR used parallel codebooks, adaptive mask supervision, and two-stage parallel decoding to improve up to 10.78% across multiple datasets, and reported +1.20% online revenue. GEAR's differentiator is treating tokenizer gradient stability as a first-class concern rather than patching it after the fact with decoding strategy. The system currently serves hundreds of millions of DAU on Douyin Ads and reports "substantial" online A/B gains, though the abstract does not disclose specific numbers.
KuaFu (Tencent) — a unified behavior compression layer whose smallest unit is a single behavior item. The problem it solves is concrete: content interest summarization tasks require reading hundreds of items per user, which serializes into tens of thousands of prompt tokens. Profiles refresh weekly across a billion users — roughly 100K QPM total. Under a fixed GPU budget, that's a hard throughput floor.
KuaFu uses a two-axis projector to compress each item to 2-4 tokens at width 128-256. That works out to roughly 10× on the token axis and 20× on the width axis, cutting per-item cache from 10 KB to 0.5 KB. Compression inevitably introduces distortion, so the paper explicitly defines four hallucination categories — fabrication, omission, date mismatch, and logical break — and pairs them with four-stage fidelity training and hierarchical intermediate evaluation. The goal is to make the compressed representation itself evaluable, rather than discovering problems only after downstream metrics degrade diffusely.
On results, KuaFu matches or exceeds uncompressed single-task production models on all five headline metrics across four production profiling tasks. Per-GPU throughput improves 37%-350%, saving 190 GPUs. On public benchmarks, it nearly sweeps existing compressors at the same compression ratio (up to +17.7 EM on out-of-domain MRQA). On RecBench, a 4B model outperforms an 8B baseline by 1.90 points. The system has run for ten months on Tencent's ads and recommendation platform, with overall GMV +1.37%. In the same space, xGR attacks from the serving side, unifying prefill and decode via separated KV caches for at least 3.49× throughput improvement. DOS uses a user-item dual-stream framework with orthogonal residual quantization for semantic IDs at Meituan. KuaFu sits closest to being a "user representation preprocessing layer" among the three.
GRP v0.1[^grp] — uses a single encoder-decoder to unify retrieval, ranking, and reward modeling, generating multimodal Semantic IDs and scoring them with a jointly trained ranking module. The frozen ranking module then provides rewards for RL post-training. The paper's mGRPO adds a reference-anchored margin to reward optimization to preserve the likelihood of logged targets — a fix for a common failure mode in generative recommendation where RL post-training skews the real online distribution. Serving optimizations cut end-to-end retrieval latency by 69%.
Online results come in two tiers: as a retrieval source alone, watch time +0.46% and shares +0.77%. With early-ranking bypass and weak-source replacement added, watch time +0.82% and shares +2.56%, with platform-level guardrail metrics neutral. The paper itself acknowledges remaining gaps in ranking quality and some recommendation metrics. This incremental deployment path is close to Gryphon, which jointly trains an item-level scoring component inside an encoder-decoder, improving Recall@1000 by 3.7% on an industrial music service and replacing 15+ candidate generators and the coarse ranking stage. DualGR uses dual-branch long-short-term routing with search-based SID decoding, achieving +0.527% video plays and +0.432% watch time in online A/B. All three report online gains in the 0.5%-3% range — evidence that generative retrieval on mature platforms is now in a "small-step replacement" phase, not a "full rewrite" phase.
Two non-deployment papers are worth noting in the roundup: Meta's Forum retrieval work transfers cross-platform Semantic IDs to a new social scenario, using a 3B instruction-tuned LLM to generate SIDs directly from user context. Spotify's VSDD models the forward corruption process of discrete diffusion as a leader-follower Stackelberg game, rewarding corruption choices by the denoiser's learning improvement rather than reconstruction difficulty.
[^grp]: The affiliations of GRP v0.1's authors are not disclosed in the abstract, so institutional attribution is omitted here.

LLM-Driven User Understanding and Context Modeling

The shared question here: under what conditions does an LLM-generated user representation actually beat a traditional aggregate profile? This week's answer is more conservative than expected.
When LLM-Inferred User Context Adds Value in Production Streaming Recommendation (production streaming platform) — a full-catalog ranking evaluation on a production streaming platform, with a 2×2 design space: representation type (aggregate or LLM-generated) × context scope (full history or attention-fused long-short-term).
The conclusion is conditional. Aggregate profiles are consistently stronger for habitual users, who make up about four-fifths of the base. LLM-generated profiles are stronger only for exploratory users — those whose subsequent interactions deviate semantically from their history. The paper also observes a popularity attractor effect in LLM profiles: intra-list diversity rises slightly, but catalog coverage drops substantially and novelty declines.
The practical implication: profile strategy should be selected by inferred consumption mode, not replaced globally. Surface-level claims that "LLM profiles are better" or "aggregate profiles are better" are both wrong. In related work, HyMiRec takes a different path — a lightweight recommender extracts coarse-grained interest embeddings, an LLM extracts fine-grained ones, and a cosine-similarity residual codebook compresses history. DMT decomposes global embedding lookup onto disjoint towers from a topological angle, achieving up to 1.9× speedup. Both avoid the assumption that the LLM alone should carry user representation.
RPTune — an end-to-end framework for SMB long-context catalog retrieval. The premise: when a merchant's catalog fits inside a long-context window, full-catalog prompting can replace the multi-stage retrieval designed for large marketplaces. But fitting isn't the same as using well — LLMs don't exploit long context uniformly.
RPTune therefore couples learned catalog curation with LLM post-training into a closed loop. An encoder-reorganizer curator ranks and prunes products based on downstream LLM feedback; the curated catalog in turn improves post-training; supervision is auto-generated by a context-relative reward, requiring no human labels. Across 7 real merchants (spanning different retail verticals), with 100 complex conversational queries per merchant, curation yields up to 31.4 percentage points of accuracy improvement, and post-training adds another 10.3 points on average, consistent across proprietary and open-weight LLMs. Worth noting: this method targets small-catalog scenarios and has not been validated on million-item marketplaces. By contrast, DARA uses RL-finetuned LLMs for few-shot budget allocation, and From Logs to Language trains a verbalization agent with RL to convert raw interaction logs into natural-language context, achieving up to 93% relative improvement over template baselines on Netflix production streaming data. RPTune's difference is making "how context is organized" itself a learnable object.
AutoResearch at Production Scale (Amazon) — extends Karpathy's AutoResearch paradigm (an LLM iteratively edits training scripts, keeping changes that improve a held-out scalar) to production scale. It ran for 12 weeks, 220+ experiments, with each iteration consuming hours of multi-GPU compute and evaluation involving multiple competing criteria.
The contribution isn't automation itself — it's the five failure modes that only emerge in production: infrastructure fragility, agent memory decay, search direction stagnation, iteration cost asymmetry, and metric fixation. The paper proposes a prevent-persist-redirect scaffolding in response. Across two independently developed representation learning systems (whose per-iteration costs differ by nearly three orders of magnitude), all five failure modes appeared simultaneously. The authors conclude this is a structural property of production-scale autonomous research, not an incidental artifact of a specific application.
On numbers, the framework achieves 1.82× Recall@6 improvement and 2.1× coherence improvement over a manual tuning baseline, and the agent autonomously designed a text-only fallback that expanded catalog coverage 5.8×. In the broader landscape, DAS uses dual-aligned semantic IDs to serve over 400M DAU in Kuaishou's ad scenario, and ScalDPP integrates determinantal point processes into RAG retrieval for inter-chunk dependency modeling. AutoResearch sits more at the "meta-layer of optimizing the process" than at any specific representation method.

Training Efficiency and Optimization for Large-Scale Recommendation

When marginal returns on model-side innovation decline, system-side waste becomes visible. Four papers this week dissect training and profiling efficiency — three from Meta.
Optimizing Effective Training Time for Large-Scale Recommendation Systems (Meta) — the paper starts with an uncomfortable number: before this work, the largest recommendation training workloads (tens of billions of training samples per day, thousands of GPUs) spent only 50%-60% of end-to-end wall-clock time actually advancing training on new data. The rest was consumed by lifecycle overhead.
The paper proposes ETT% (Effective Training Time) as an operational framework, hierarchically attributing lost time to independent infrastructure components (TTS/TTR/failure counts, etc.) and exposing work repeated across job restarts. Based on this attribution, the team implemented a set of full-stack optimizations: communication elimination and pipeline overlap during trainer initialization, dynamic shape handling, autotuning pruning, reusable PyTorch 2 compilation caches, asynchronous checkpointing, independent model publishing, and reduced recovery cost.
Results: ETT% improved on every benchmark, averaging +15.5%, with the largest workload reaching 85%. After deployment, fleet-wide ETT% rose from around 80% to above 90%. The significance of this metric is that it expands "training efficiency" from a single MFU number into a decomposable operational language. For reference, Kunlun raised MFU on B200 from 17% to 37% and doubled scaling efficiency through unified architecture design, while xGR extracts throughput from the serving side. All three target different facets of the same class of waste.
Component Benchmark (Meta) — a hierarchical profiling system targeting the structural heterogeneity of recommendation models: memory-bandwidth-bound operators, small compute-bound dense layers, dynamic shapes from jagged categorical features, and low-arithmetic-intensity operators all mixed in one graph. Standard tools give only end-to-end throughput or operator-level traces, with no way to attribute to the submodules modeling engineers actually discuss.
CB profiles each submodule independently and outputs tree-structured interactive visualizations, with a pluggable submodule benchmark framework underneath. The target object model is terabyte-scale, runs on thousands of GPUs, and ingests 100B samples per day. The paper gives no quantified speedup — it's a tooling paper — but it complements DMT's tower decomposition: DMT trades structural splitting for data locality, CB trades profiling granularity for attribution capability.
UWSR — explains the widespread one-epoch phenomenon in recommendation models through violations of the prequential principle. In the first epoch, a sample's label has not yet influenced the embedding rows used to score it. In later epochs, those rows already contain shifts introduced by prior updates, giving the shared consumer an incentive to exploit that shift — which doesn't generalize. The paper calls this self-influence asymmetry and connects it to exact scalar models and local influence analysis.
Based on this hypothesis, the paper proposes uncertainty-weighted sensitivity regularization (UWSR), which penalizes the consumer's reliance on uncertain embeddings to counteract the mismatch. The key difference from other remedies: UWSR keeps learned embeddings rather than resetting them. Across three benchmarks, four-epoch UWSR reduces test cross-entropy by 1.38%-6.78% and improves AUC by 0.0058-0.0231 relative to single-epoch training. This perspective differs from the sparsification direction in Sparse Feature Attention, but both fall under "finding gains in training dynamics rather than structural design."
Soft Curriculum Learning (short-video platform) — the goal is breaking the popularity feedback loop. Training data on short-video platforms is naturally skewed toward head items; models memorize head patterns at the expense of long-tail generalization. Curriculum learning is the obvious remedy, but the industrial obstacle is that dynamic data rejection algorithms are heavily CPU-bound and starve TPUs/GPUs.
The paper's solution avoids rigid data filtering, instead using loss annealing plus reweighting within the graph to achieve dynamic curriculum pacing in continuous training without sacrificing system throughput. Validation covers three model types: sequence-based retrieval (e.g., SASRec), two-tower retrieval, and large-scale continuous ranking models. Online A/B shows improvements in overall user satisfaction and fresh content consumption, with specific magnitudes undisclosed and no throughput degradation. In historical context, MEMOIR uses LLMs to segment interaction history into semantic memories to handle preference drift, and Factorized Transport uses optimal transport approximation to unify multimodal multi-view representations. Both focus on representation; soft curriculum learning focuses on distribution control of training data.

Industrial Retrieval and Ranking Architecture Optimization

All four papers come from production systems. The common theme: structural modifications to existing architectures, not new paradigms.
HELIX (ByteDance) — the starting point is an observation: industrial ranking models scale along two axes, heterogeneous feature interaction and long behavior sequence modeling. Scaling either axis alone hits a ceiling and a suboptimal scaling-law slope. The paper hypothesizes that better slopes require scaling both axes jointly.
HELIX interleaves sequence retrieval with feature interaction and enforces unidirectional information flow from reusable sequence states to candidate-conditioned mix-tokens. This preserves cross-depth communication while making user-side sequence computation amortizable, enabling asymmetric flexible scaling of sequence modeling and feature interaction. After deployment on TikTok e-commerce recommendation, offline CTR AUC, CVR AUC, and other ranking metrics improved consistently. In online A/B, per-capita GMV on e-commerce videos rose about 6%.
This result forms a line with prior work in the same space. HyFormer already integrated long sequence modeling and feature interaction into a single hybrid Transformer backbone, outperforming LONGER and RankMixer on industrial data. IAT uses a two-stage method to compress each historical interaction instance into a unified instance token. LMN uses external memory modules for long-term interest modeling in Douyin e-commerce search. HELIX's difference is elevating "how the two axes should couple" into an explicit architectural constraint, with unidirectional information flow as the concrete form.
ESP (LinkedIn) — targets a structural limitation of two-tower retrieval: it compresses heterogeneous signals — semantic relevance, engagement, revenue — into a single static embedding space, with no way to adjust objective priorities at serving time after training. And joint multi-objective loss optimization often causes inter-objective interference.
Embedding Subspace Partitioning decomposes the embedding into task-aware subspaces, replacing the single inner product with a weighted sum of subspace similarities, with weights adjustable at serving time. For Transformer two-towers, ESP uses the model's native EOS token as a segmenter, paired with segment-aware attention masking and positional encoding resets, guaranteeing subspace isolation in a single forward pass. On the serving side, GPU-accelerated exhaustive kNN runs on a single concatenated index, avoiding the per-objective ANN infrastructure that multi-head methods require.
On an open benchmark built from MS MARCO, a single ESP model traces a wide Pareto frontier, consistently outperforming multi-task baselines across operating points. On LinkedIn's job matching platform (70M+ weekly active users), ESP enabled dynamic retrieval reconfiguration and improved key business metrics. In related history, DSIN segments behavior sequences into 30-minute sessions to model interest evolution, RAT enhances candidate representations by retrieving external similar samples, and HiGR does multi-objective preference alignment in list-level generation. Both ESP and HiGR handle multiple objectives, but ESP puts the adjustment knob at serving time.
Retail Product Search at Target (Target) — a complete engineering record of a hybrid retrieval system fusing lexical and vector retrieval. The paper covers data processing, embedding training, result-set precision control, multi-channel fusion, and low-latency optimization. For fusion, the team compared multiple approaches and adopted weighted interleaving.
Online A/B against pure lexical retrieval: CTR +0.97%, order conversion +0.98%, per-capita demand +1.10%, and zero-result searches roughly halved. The system is deployed at scale, serving millions of users daily. The value of this work is mainly engineering completeness — retail search intent spans from exact match to open-ended discovery, while also balancing relevance, revenue, and margin. No single method covers it.
TSG Suggester (AWS) — troubleshooting guide (TSG) recommendation for cloud incident management. On-call engineers manually keyword-search guides under time pressure; prior empirical research found guide retrieval consumes a substantial fraction of mitigation time.
The paper compares five retrieval strategies on 314 real incidents (spanning 112 TSGs and 18 service teams). Tree+KG converts each guide into a tree preserving native section hierarchy, mounting LLM-generated question abstractions at internal nodes — the goal is bridging the "solution-oriented" language of guides with the "problem-oriented" language of incidents — then extracts a per-guide entity knowledge graph, fusing embedding similarity with entity-level matching at query time.
Results: Top-1 accuracy 54.78%, Top-5 82.48%, leading all baselines at every cutoff, with an 8.58-point Top-1 improvement over pure-text RAG. Two secondary findings are more notable. Structural alignment dominates: methods that preserve or reconstruct document structure clearly outperform flat chunking when precise discrimination is needed. Multimodal augmentation actively hurts: generating captions for guide screenshots and injecting them loses 22.64 Top-5 points relative to the pure-text baseline, because generic captions dilute the embedding rather than sharpening it. This negative result is rare in RAG work. In related directions, RAGSearch's benchmark work finds that agentic search substantially improves dense RAG and narrows the gap with GraphRAG, while GraphRAG retains an edge in complex multi-hop reasoning. Mandol uses a hierarchical memory model with agglomerative semantic structure for 5.4× retrieval speedup. TSG Suggester's results partially support the structure-first judgment.

Directions to Watch

Tokenizer stability is becoming the engineering focal point of generative recommendation. GEAR reparameterizes the codebook with orthogonal bases to handle representation collapse. KuaFu uses a two-axis projector plus fidelity training to handle compression distortion. MDGR uses parallel codebooks and adaptive masking. What these share: the tokenizer is no longer a preprocessing step but a training object jointly optimized with the generator. The evidence: GEAR explicitly states that expanding the codebook to ease collisions worsens collapse — meaning the two cannot be tuned separately. The drivers are mainly teams like ByteDance and Tencent with billion-scale candidate pools, since collision problems aren't significant at small-to-medium scale.
Behavior compression layers may become a standalone component in the LLM recommendation stack. KuaFu's shape is worth noting: it produces no recommendations, only reusable compressed representations. Per-item cache drops from 10 KB to 0.5 KB, saving 190 GPUs, with GMV +1.37%. This "representation layer, not model" positioning means it can be shared across multiple downstream tasks rather than bound to one model. Tencent has run it for ten months, demonstrating engineering viability. Practical adoption depends on whether the compression ratio and fidelity hold across more task types — KuaFu currently validates four profiling tasks.
How training efficiency is measured is shifting from a single metric to an attributable framework. Meta's ETT% decomposes the 50%-60% effective training time down to specific infrastructure components. Kunlun uses MFU and scaling efficiency. Component Benchmark uses submodule-level hierarchical profiling. The value of this body of work to the industry may not lie in any single optimization, but in establishing a comparable, locatable language. At thousands of GPUs and tens of billions of samples per day, "where is it slow" is harder to answer than "how do we make it fast."

Paper Roundup

Generative Recommendation
GEAR — ByteDance proposes an end-to-end generative ad retrieval framework, using orthogonal-basis codebook reparameterization (BasisVQ/BasisRQ) to mitigate representation collapse and a context-conditioned reranking head to resolve item collisions; serves hundreds of millions of DAU on Douyin Ads.
KuaFu — Tencent proposes a unified behavior compression layer, compressing each behavior to 2-4 tokens at width 128-256, cutting cache from 10 KB to 0.5 KB; per-GPU throughput improves 37%-350%, saving 190 GPUs, with online GMV +1.37%.
GRP v0.1 — Uses a single encoder-decoder to unify retrieval, ranking, and reward modeling, proposing mGRPO to preserve logged-target likelihood; serving latency reduced 69%, online watch time +0.46%~+0.82% and shares +0.77%~+2.56%.
Exploring Forum Post Retrieval with Generative Modeling — Meta reuses hierarchical SIDs from cross-platform Facebook Feed in the new Facebook Forum scenario, using a 3B instruction-tuned LLM to generate SIDs, with ablations on SID construction, user history, and profile features.
VSDD — Spotify models the forward corruption process of discrete diffusion as a leader-follower Stackelberg game, rewarding corruption choices by the denoiser's learning improvement rather than reconstruction difficulty; improvements across molecular validity, text perplexity, and offline playlist recommendation metrics.
LLM User Understanding and Context Modeling
When LLM-Inferred User Context Adds Value — A 2×2 controlled experiment on a production streaming platform finds aggregate profiles stronger for about four-fifths of habitual users, LLM-generated profiles better only for exploratory users, with a popularity attractor effect present.
RPTune — Couples learned catalog curation with LLM post-training, using context-relative reward to auto-generate supervision; across 7 real merchants, curation yields up to 31.4 points of improvement, with post-training adding another 10.3 points on average.
AutoResearch at Production Scale — Amazon runs 12 weeks and 220+ experiments, distilling five failure modes of production-scale autonomous research and proposing a prevent-persist-redirect scaffolding; Recall@6 improves 1.82×, coherence 2.1×, and catalog coverage 5.8×.
Training Efficiency and Optimization
Optimizing Effective Training Time — Meta proposes the ETT% operational framework, attributing training lifecycle overhead to specific infrastructure components; benchmark ETT% +15.5% on average, fleet-wide from around 80% to above 90%.
Component Benchmark — Meta proposes a submodule-level hierarchical profiling system with a pluggable architecture and tree-structured interactive visualization, targeting recommendation models at terabyte scale, thousands of GPUs, and 100B samples per day.
UWSR — Explains the one-epoch phenomenon via prequential principle violations and self-influence asymmetry, proposing uncertainty-weighted sensitivity regularization; four-epoch training reduces test cross-entropy by 1.38%-6.78% and improves AUC by 0.0058-0.0231 relative to single-epoch.
Soft Curriculum Learning — A short-video platform replaces rigid data filtering with loss annealing and in-graph reweighting, breaking the popularity feedback loop in continuous training across sequence retrieval, two-tower retrieval, and large-scale ranking, with online user satisfaction and fresh content consumption improving and no throughput degradation.
Industrial Retrieval and Ranking Architecture
HELIX — ByteDance proposes a unified ranking architecture interleaving sequence retrieval with feature interaction, enforcing unidirectional information flow from reusable sequence states to candidate-conditioned mix-tokens; deployed on TikTok e-commerce, with online per-capita GMV up about 6%.
ESP — LinkedIn decomposes two-tower embeddings into task-aware subspaces, replacing the single inner product with a serving-time-adjustable weighted similarity and using EOS token segmentation to guarantee subspace isolation; enables dynamic retrieval reconfiguration on a job matching platform with 70M+ weekly active users.
Retail Product Search at Target — Target's hybrid retrieval system fuses lexical and vector retrieval with a weighted interleaving fusion strategy; online CTR +0.97%, order conversion +0.98%, per-capita demand +1.10%, and zero-result searches halved.
TSG Suggester — AWS proposes a Tree+KG retrieval system, preserving guide section hierarchy and mounting LLM-generated question abstractions at internal nodes; on 314 real incidents, Top-1 54.78% and Top-5 82.48%, 8.58 Top-1 points above pure-text RAG, and reports that multimodal caption injection actually loses 22.64 Top-5 points.
  • Recommendation Systems
  • Weekly
  • Papers
  • AI Weekly 2026-W40AI Tech Daily - 2026-10-03
    Loading...