The short answer
Three things to know
- 1Before their first birthday, infants reorganise their perception of sound around the of their native language, making some foreign contrasts genuinely harder to hear as distinct.
- 2Adult brains are not locked: targeted training — especially practice with many different speakers and contexts — can shift category boundaries and improve discrimination of non-native sounds.
- 3The difficulty is not a failure of intelligence or effort; it is the predictable cost of a system that became very good at one language very early.
01 · The raw material
Sound arrives as a continuum
The physical world of speech is not pre-sliced. Acoustic energy varies continuously along dimensions such as — the gap between a consonant release and the start of vocal-fold vibration — and formant frequency, the resonant peaks that distinguish vowels. There are no natural gaps in this stream that announce where one phoneme ends and another begins. Every language must impose its own set of boundaries on this shared acoustic space, and different languages draw those lines in different places.
A speaker of one language and a speaker of another are, in a real sense, hearing the same raw signal through different perceptual grids. The grids are not innate; they are learned. What makes the learning remarkable is how early and how thoroughly it happens, and how durable the result turns out to be across a lifetime of subsequent experience.
02 · The infant sorting machine
Categories form before the first word
Research synthesised by Patricia Kuhl describes a process in which infants begin as what she calls universal listeners, capable of discriminating phonetic contrasts from any of the world's languages. Over the first months of life, exposure to the surrounding language causes perception to reorganise around the sound categories of that language. Sounds cluster toward prototypical examples — a phenomenon Kuhl terms the — so that within-category variation becomes harder to detect while between-category differences become more salient.
By roughly the end of the first year, sensitivity to many non-native contrasts has already declined measurably. The infant has not lost hearing acuity; it has gained a highly efficient native-language filter. That filter is the foundation of fluent perception in the mother tongue, and it is also the source of the difficulty that awaits any adult who later tries to learn a language whose phoneme boundaries fall in unfamiliar places.
Continuous acoustic space
Speech energy varies along dimensions such as voice-onset time and formant frequency without natural breaks. No language boundary exists yet in the infant's perception.
All contrasts potentially discriminableNative-language categories form
Exposure to the surrounding language causes perception to cluster around prototypical native-language sounds. Within-category variation becomes less salient; between-category differences become more so.
Reorganisation largely complete by ~12 monthsSecond-language boundaries fall elsewhere
A second language may draw its phoneme boundaries at different points on the same acoustic axes. Two distinct second-language sounds may land inside a single native-language category box.
Overlap zone: where confusion is most likelyTraining shifts the boundary
Deliberate perceptual training — especially with high acoustic variability — can move the effective category boundary, allowing the learner to hear the second-language contrast as distinct.
Generalisation requires varied-speaker exposure03 · The collision of categories
When two sounds share one box
The practical consequence for adult learners is that two phonemes in a second language may map onto a single category in the native language. A well-studied example involves Japanese and English: the English contrast between the sounds represented by the letters R and L corresponds to a voice-onset-time and formant region that Japanese phonology treats as a single category. A native Japanese listener is not mishearing; the perceptual system is doing exactly what it was trained to do — collapsing variation within a category rather than treating it as meaningful.
Melissa Baese-Berk and colleagues reviewing the literature on non-native speech sound representations describe how these native-language representations actively compete with the formation of new second-language categories. The problem is not simply that the new contrast is unfamiliar; it is that an existing representation is already occupying the relevant perceptual territory, and that representation has years of reinforcement behind it. Overcoming it requires more than exposure — it requires the perceptual system to build a genuinely new boundary.
04 · Plasticity persists
The adult brain can still move its lines
The older view — that a critical period closes and adult phonetic learning becomes essentially impossible — is not supported by the current evidence. Adults do show reduced plasticity compared with infants, and the reorganisation that happens effortlessly in the first year requires deliberate effort later. But plasticity is not absent. Laboratory training studies consistently show that adults can learn to discriminate non-native contrasts they initially could not hear as distinct, and that this learning reflects genuine perceptual change rather than a post-perceptual decision strategy.
The neural commitment model, as Kuhl frames it, holds that early language experience commits neural circuits to native-language patterns in ways that make subsequent reorganisation harder but not impossible. This framing is more precise than a simple critical-period story: it predicts that difficulty will vary with how much the second-language category overlaps with existing native-language representations, which is broadly what the empirical record shows. Contrasts that fall entirely outside native-language category space are often easier for adults to acquire than those that fall inside an existing category.
05 · The training evidence
Variability is the key ingredient
A primary study by Lim and Holt examined high-variability versus low-variability phonetic training and found that exposure to many different talkers and acoustic contexts produced learning that generalised more broadly than training on a narrow, consistent stimulus set. This matters because the goal of language learning is not to recognise one speaker's version of a sound in one recording condition; it is to recognise that sound across the full range of speakers, speaking rates, and environments a learner will actually encounter.
The mechanism proposed is that variability forces the perceptual system to extract the abstract category rather than memorise specific acoustic tokens. When training is too uniform, learners may improve on the trained items without building a representation robust enough to transfer. High-variability training is more demanding and initially produces slower apparent progress, but the resulting representations appear more durable and more general — a finding with direct implications for how language instruction might be designed.
Before their first birthday, infants reorganise their perception of sound around the phoneme categories of their native language, making some foreign contrasts genuinely harder to hear as distinct.
Most phonetic training studies are short-term laboratory experiments. Whether the perceptual improvements they demonstrate translate into the fast, automatic, noise-robust perception of a fluent speaker remains an open and important question. Individual differences in learning rate are large and poorly understood. The link between better phoneme discrimination and better communicative ability in a second language is plausible but not yet tightly established by the available evidence.
06 · What this means in practice
Difficulty is structural, not personal
Understanding the perceptual basis of phonetic difficulty reframes what it means to struggle with foreign sounds. The adult learner who cannot reliably hear the difference between two second-language phonemes is not inattentive or untalented; they are experiencing the predictable consequence of a perceptual system that was optimised, very successfully, for a different language. The difficulty is structural, and it calls for structural solutions: deliberate perceptual training, not just more conversation, and exposure designed to build generalisable categories rather than familiarity with a single accent.
Open questions remain. Most training studies are conducted in laboratory conditions over short periods, and how well laboratory gains translate into the kind of fluent, automatic perception that characterises native speakers is not yet clear. Individual differences in learning rate are large and not fully explained. And the relationship between improved phonetic discrimination and improved communicative ability in a second language — while plausible — is not as tightly established as researchers would like. The field is active, and the practical payoff of this basic science is still being worked out.
07 · Sources
Evidence behind this article
This article draws on three peer-reviewed sources. Confidence ratings reflect the type and recency of each source, not the importance of its findings.
- 01Kuhl · A new view of language acquisitionPeer-reviewed review ↗
Patricia Kuhl's review in the Proceedings of the National Academy of Sciences synthesises evidence for the perceptual magnet effect and the native-language neural commitment model, arguing that early language experience reorganises infant speech perception in ways that shape — and constrain — later learning.
- 02Baese-Berk · The nature of non-native speech sound representationsPeer-reviewed review ↗
Melissa Baese-Berk and colleagues review the literature on how non-native speech sounds are represented in the adult mind, examining why existing native-language categories interfere with the formation of new second-language categories and what conditions support change.
- 03High- and low-variability phonetic trainingPrimary study ↗
This primary study compares high-variability and low-variability phonetic training regimens, finding that exposure to a wider range of talkers and acoustic contexts produces learning that generalises more broadly to untrained speakers and conditions.
