Research

Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers

arXiv:2601.15014v2 Announce Type: replace-cross Abstract: We study in-context learning for nonparametric regression with alpha-Holder smooth regression functions, for some alpha>0. We prove that, with

DGX agentpaper
researcharxiv-cs-lg

arXiv:2601.15014v2 Announce Type: replace-cross Abstract: We study in-context learning for nonparametric regression with alpha-Holder smooth regression functions, for some alpha>0. We prove that, with n in-context examples and d-dimensional regression covariates, a pretrained transformer with Theta(log n) parameters and Omegaigl(n^{2alpha/(2alpha+d)}log^3 nigr) pretraining sequences can achieve the minimax optimal rate of convergence Oigl(n^{-2alpha/(2alpha+d)}igr) in mean squared error. Our result requires substantially fewer transformer parameters and pretraining sequences than previous results in the literature. This is achieved by showing that transformers are able to approximate local polynomial estimators efficiently by implementing a kernel-weighted polynomial basis and then running gradient descent.

Source: arXiv cs.LG | 2026-05-20

Loading related sources…