We study speaker-disjoint accent identification for L2 English, where the goal is to predict a speaker’s firstlanguage (L1) background from English pronunciation. Most existing systems classify accents using a single utterance-level representation, but such global representations can obscure pronunciation cues that depend on specific English phonemes. We propose a transcript-assisted model that makes phoneme information explicit during accent identification. Instead of representing an utterance only as a global speech embedding, we represent it as a sequence of pronunciation units, each combining acoustic evidence from a spoken segment with the aligned English phoneme for that segment. A frozen speech encoder provides the acoustic features, while the transcript is used only to obtain phoneme-level forced alignments. No word-level or sentence-level text representation is passed to the accent classifier. Under a four-fold speaker-disjoint protocol on L2-ARCTIC, our model achieves 81.41% accuracy and 81.21% macro-F1, the highest mean performance among the evaluated systems. Diagnostic ablations support the importance of phoneme-aligned token construction, while a Whisper-based ablation shows an additional gain from phoneme information.
Phoneme-aware pronunciation representations for L2-english L1-background accent identification
Submitted to ArXiV, 9 September 2026
Type:
Rapport
Date:
2026-09-09
Department:
Sécurité numérique
Eurecom Ref:
8964
Copyright:
© EURECOM. Personal use of this material is permitted. The definitive version of this paper was published in Submitted to ArXiV, 9 September 2026 and is available at :
See also:
PERMALINK : https://www.eurecom.fr/publication/8964