Modeling and recognition of phonetic and prosodic factors for improvements to acoustic speech recognition models

Borys, Sarah; Cohen, Aaron; Hasegawa-Johnson, Mark; Cole, Jennifer

doi:10.21437/Interspeech.2004-756

This paper examines the usefulness of including prosodic and phonetic context information in the phoneme model of a speech recognizer. This is done by creating a series of prosodic and phonetic models and then comparing the mutual information between the observations and each possible context variable. Prosodic variables show improvement less often than phone context variables, however, prosodic variables generally show a larger increase in mutual information. A recognizer with allophones defined using the maximum mutual information prosodic and phonetic variables outperforms a recognizer with allophones defined exclusively using phonetic variables.

Modeling and recognition of phonetic and prosodic factors for improvements to acoustic speech recognition models

Sarah Borys, Aaron Cohen, Mark Hasegawa-Johnson, Jennifer Cole