Model Releases
From Scaling to Structured Expressivity: Rethinking Transformers for CTR Prediction
arXiv:2511.12081v2 Announce Type: replace-cross Abstract: Despite massive investments in scale, deep models for click-through rate (CTR) prediction often exhibit rapidly diminishing returns -- a stark
arXiv:2511.12081v2 Announce Type: replace-cross Abstract: Despite massive investments in scale, deep models for click-through rate (CTR) prediction often exhibit rapidly diminishing returns -- a stark contrast to the {predictable scaling laws} seen in large language models (LLMs). We identify the root cause as a {fundamental} extit{structural misalignment}: {standard} Transformers assume sequential compositionality, whereas CTR data demand combinatorial reasoning over {heterogeneous} fields. To restore alignment, we introduce the extbf{Field-Aware Transformer (FAT)}. {By reconstructing the standard Transformer block with field-centric parameters, FAT achieves extit{structured expressivity}, {fundamentally shifting the model complexity dependence from the total vocabulary size n with the number of fields F (n gg F).}} Crucially, to decouple model capacity from field cardinality, FAT employs a {{Basis-Composed Hypernetwork}} to synthesize field-specific parameters from shared bases, further reducing parameter complexity. {Theoretically, we ground this scaling behavior through a formal scaling law based on Rademacher complexity. Empirically, FAT outperforms exisiting state-of-the-art methods with up to extbf{{+4.38%}} AUC improvement, and delivers extbf{+2.33%} CTR and extbf{+0.66%} RPM in live production.} Our work establishes that scalable recommendation arises not from size alone, but from extit{structured expressivity} -- architectural coherence with data semantics.
Source: arXiv cs.LG | 2026-06-02