Details

Title

An Effective Speaker Clustering Method using UBMand Ultra-Short Training Utterances

Journal title

Archives of Acoustics

Yearbook

2016

Volume

vol. 41

Numer

No 1

Publication authors

Keywords

automatic speech recognition; interindividual difference compensation; speaker clustering; universal background model; GMM weighting factor adaptation

Divisions of PAS

Nauki Techniczne

Description

Archives of Acoustics is an English-language peer-reviewed quarterly journal publishing original research papers from all areas of acoustics and abstracts from some specialised acoustical conferences. It gives free internet access to its full content (abstracts of research papers) to current issues.

Archives of Acoustics, the peer-reviewed quarterly journal publishes original research papers from all areas of acoustics like:

  • acoustical measurements and instrumentation,
  • acoustics of musics,
  • acousto-optics,
  • architectural, building and environmental acoustics,
  • bioacoustics,
  • electroacoustics,
  • linear and nonlinear acoustics,
  • noise and vibration,
  • physical and chemical effects of sound,
  • physiological acoustics,
  • psychoacoustics,
  • quantum acoustics,
  • speech processing and communication systems,
  • speech production and perception,
  • transducers,
  • ultrasonics,
  • underwater acoustics.

Earlier issues are available on the old website http://acoustics.ippt.gov.pl/index.php/aa/issue/archive

Abstract

The same speech sounds (phones) produced by different speakers can sometimes exhibit significant differences. Therefore, it is essential to use algorithms compensating these differences in ASR systems. Speaker clustering is an attractive solution to the compensation problem, as it does not require long utterances or high computational effort at the recognition stage. The report proposes a clustering method based solely on adaptation of UBM model weights. This solution has turned out to be effective even when using a very short utterance. The obtained improvement of frame recognition quality measured by means of frame error rate is over 5%. It is noteworthy that this improvement concerns all vowels, even though the clustering discussed in this report was based only on the phoneme a. This indicates a strong correlation between the articulation of different vowels, which is probably related to the size of the vocal tract.

Publisher

Committee on Acoustics PAS, PAS Institute of Fundamental Technological Research, Polish Acoustical Society

Identifier

ISSN 0137-5075 ; eISSN 2300-262X

DOI

10.1515/aoa-2016-0011

×