Tutorials
LLMs learn backwards, and the scaling hypothesis is bounded. [D]
This r/MachineLearning discussion post argues that LLMs acquire knowledge in a counterintuitive 'backwards' order during training — learning complex, high-level patterns before simpler foundational on
This r/MachineLearning discussion post argues that LLMs acquire knowledge in a counterintuitive "backwards" order during training — learning complex, high-level patterns before simpler foundational ones — and uses this framing to challenge the indefinite validity of the scaling hypothesis. The post contends that because LLMs face irreducible theoretical limits (such as the fact that the test loss of an LLM is always irreducible and cannot be brought below a theoretical bound, irrespective of model or pre-training data size ), continued scaling of compute and parameters will yield diminishing returns. The thread likely draws on growing community skepticism around whether the continuation of scaling has recently been called into question, raising the concern of whether scaling will hit a wall and what other paths forward exist .
Related
- What do Language Models Learn and When? The Implicit Curriculum Hypothesis
- Learning is Forgetting: LLM Training As Lossy Compression
- How to sketch a learning algorithm
- A Severity-Based Curriculum Learning Strategy for Arabic Medical Text Generation
Source: r/MachineLearning | 2026-04-12