Model Releases

Geometry-Constrained Kolmogorov-Arnold Networks: Learning Edge Geometry via Banach Duality

arXiv:2608.25807v1 Announce Type: new Abstract: Kolmogorov-Arnold Networks (KANs) replace fixed activations in deep architectures with learnable univariate edge functions, making the choice of edge pa

DGX agentpaper
model-releasesarxiv-cs-lg

arXiv:2608.25807v1 Announce Type: new Abstract: Kolmogorov-Arnold Networks (KANs) replace fixed activations in deep architectures with learnable univariate edge functions, making the choice of edge parametrisation central. Existing variants rely on fixed bases such as splines, polynomials, or Fourier features, which impose a function-space geometry before data are observed. We introduce geometry-constrained KANs, a family of edge activations derived from Banach duality maps in which the geometry itself is learned through a scalar exponent p > 1 per edge. This exponent controls the qualitative response: sub-Euclidean values produce sharp, threshold-like behaviour reminiscent of the ell_1 (LASSO) geometry, p = 2 recovers the linear regime, and larger values produce flatter responses near the origin. Across 50 symbolic-regression targets (40 from the AI Feynman benchmark plus 10 synthetic stress tests), geometry-constrained KANs match or beat every fixed-basis baseline on median NRMSE (Banach-KAN 0.030, tying Chebyshev and improving on splines); on average rank Banach-KAN is best on the 18-equation core (2.00) and statistically tied with the strongest spline on the full benchmark (2.32 vs. 2.34). The clearest gains appear under measurement noise: as sigma grows from 0 to 1, ell^p-KAN degrades only 3.7imes -- below even a cross-validated spline (approx 11imes) -- while an unregularised spline degrades 21.6imes; Banach-KAN degrades 8.8imes, comparable to a tuned spline but far more stable than the unregularised one. Banach-KAN also takes the most per-equation wins in the small-sample regime, with fixed-basis models catching up only as the training set grows. Learned exponents provide an interpretable, relative signal: at a fixed initialisation they reveal a consistent, target-dependent geometric ordering across equation families and input dimensions.

Related

Source: arXiv cs.LG | 2026-08-27

Loading related sources…