Applications

Approximation Error Upper and Lower Bounds for Holder Class with Transformers

arXiv:2605.07463v1 Announce Type: new Abstract: We explore the expressive power of Transformers by establishing precise approximation error upper and lower bounds for Holder class. Specifically, a new

DGX agentpaper
applicationsarxiv-cs-lg

arXiv:2605.07463v1 Announce Type: new Abstract: We explore the expressive power of Transformers by establishing precise approximation error upper and lower bounds for Holder class. Specifically, a new approximation upper bound is derived for the standard Transformer architecture equipped with Softmax operators, ReLU activation functions, and residual connections. We prove that a Transformer network composed of at most O(arepsilon^{-{d_{0}}/{alpha}}) blocks can approximate any bounded Holder function with d_{0}-dimensional input and smoothness alphain(0,1] under any accuracy arepsilon>0. In the case of approximation lower bounds, leveraging the VC-dimension upper bound, we are the first to rigorously prove that Transformers demand for at least Omega(arepsilon^{-{d_{0}}/({4alpha})}) blocks to achieve the arepsilon approximation accuracy. As a final step, we extend the derived results for standard Transformers to a general regression task and establish the corresponding excess risk rates demonstrating Transformers' empirical effectiveness in real-world settings.

Source: arXiv cs.LG | 2026-05-11

Loading related sources…