Ecole d'ingénieur et centre de recherche en télécommunications

The LIA-Eurecom RT'09 speaker diarization system : enhancements in speaker modelling and cluster purification

Bozonnet, Simon; Evans, Nicholas; Fredouille, Corinne

ICASSP 2010, 35th International Conference on Acoustics, Speech, and Signal Processing, March 14-19, 2010, Dallas, Texas, USA

There are two approaches to speaker diarization. They are bottom-up and top-down. Our work on top-down systems show that they can deliver competitive results compared to bottom-up systems and that they are extremely computationally efficient, but also that they are particularly prone to poor model initialisation and cluster impurities. In this paper we present enhancements to our state-of-the-art, top-down approach to speaker diarization that deliver improved stability across three different datasets composed of conference meetings from five standard NIST RT evaluations. We report an improved approach to speaker modelling which, despite having greater chances for cluster impurities, delivers a 35% relative improvement in DER for the MDM condition. We also describe new work to incorporate cluster purification into a top-down system which delivers relative improvements of 44% over the baseline system without compromising computational efficiency.

Document Doi Hal Bibtex

Mots Clés:Speaker diarization, speaker segmentation, speaker clustering, cluster purification, DER, MDM, SDM
Département:Communications Multimédia
Eurecom ref:3000
Copyright: © 2010 IEEE. Personal use of this material is permitted. However, permission to reprint/republish this material for advertising or promotional purposes or for creating new collective works for resale or redistribution to servers or lists, or to reuse any copyrighted component of this work in other works must be obtained from the IEEE.
Bibtex: @inproceedings{EURECOM+3000, doi = { }, year = {2010}, title = {{T}he {LIA}-{E}urecom {RT}'09 speaker diarization system : enhancements in speaker modelling and cluster purification}, author = {{B}ozonnet, {S}imon and {E}vans, {N}icholas and {F}redouille, {C}orinne }, booktitle = {{ICASSP} 2010, 35th {I}nternational {C}onference on {A}coustics, {S}peech, and {S}ignal {P}rocessing, {M}arch 14-19, 2010, {D}allas, {T}exas, {USA}}, address = {{D}allas, {\'{E}}{TATS}-{UNIS}}, month = {03}, url = {} }
Voir aussi: