Model Releases
The Learning Objective Governs Perceptual Narrowing: A Cross-Lingual, Layer-Wise, Ten-Seed Study of Self-Supervised Speech Encoders
arXiv:2608.00507v1 Announce Type: new Abstract: Perceptual narrowing---the developmental loss of non-native phoneme discrimination in the first year of life itep{werker1984}---is a canonical developme
arXiv:2608.00507v1 Announce Type: new Abstract: Perceptual narrowing---the developmental loss of non-native phoneme discrimination in the first year of life itep{werker1984}---is a canonical developmental finding, yet what learning objective produces it remains open. We train a (sim)7,M-parameter Transformer encoder on child-directed and read speech and evaluate phoneme ABX in English, French, and Mandarin over ten seeds, the seed as the unit of replication. Six results. extbf{(1)}~The objective sets the direction of cross-lingual transfer: reconstruction (masked mel-prediction) degrades non-native discrimination, prediction (frame-contrastive) improves it---a same-encoder, same-data gap of (+0.051) in first-layer Mandarin ABX ((p=3imes10^{-8})), unanimous in sign across twenty runs. extbf{(2)}~That decline combines a large arm-intrinsic difficulty gradient with a smaller language-specialization effect (matched vs. mismatched (+0.022), (p=10^{-4}), all four layers). extbf{(3)}~Against a language-symmetric raw-mel floor, reconstruction pushes the first layer below the discriminability of its input; prediction pushes it above. extbf{(4)}~Read speech gives a (3.6imes) steeper non-native decline than child-directed speech. extbf{(5)}~The customary three-seed budget cannot see this reliably: an effect unambiguous at ten seeds is called significant by as few as 70% of three-seed subsets. extbf{(6)}~Six objective configurations---sharpening, compression, consolidation, their composition, and word-level semantic grounding in two forms---fail to produce the full developmental signature (native improves and non-native declines): a single objective moves both languages the same way because it acts on a shared representation. We conclude that the objective, not the architecture, is the first-order determinant of narrowing-shaped representational change.
Related
- How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures
- Analyzing Speech Condition Effects in Dysarthric ASR: A Layer-wise Probing Study
- Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
Source: arXiv cs.CL | 2026-08-04