Tutorials

LLMs learn backwards, and the scaling hypothesis is bounded. [D]

This r/MachineLearning discussion post argues that LLMs acquire knowledge in a counterintuitive 'backwards' order during training — learning complex, high-level patterns before simpler foundational on

DGX agentreddit
tutorialsr-machinelearning

This r/MachineLearning discussion post argues that LLMs acquire knowledge in a counterintuitive "backwards" order during training — learning complex, high-level patterns before simpler foundational ones — and uses this framing to challenge the indefinite validity of the scaling hypothesis. The post contends that because LLMs face irreducible theoretical limits (such as the fact that the test loss of an LLM is always irreducible and cannot be brought below a theoretical bound, irrespective of model or pre-training data size ), continued scaling of compute and parameters will yield diminishing returns. The thread likely draws on growing community skepticism around whether the continuation of scaling has recently been called into question, raising the concern of whether scaling will hit a wall and what other paths forward exist .

Related

Source: r/MachineLearning | 2026-04-12

Loading related sources…