Autocorrection of Speech Recognition Errors Using Confusion Network GUI


University of Toronto (Jan 2017 - Jan 2020)


Project Summary

  • Developed a graphical web interface in Node.JS that shows a visual representation of a speech recognition confusion network, from a lattice, that can be manipulated through touch in order to correct speech recognition errors (based on “Parakeet” by Vertanen et. al [IUI ‘09])
  • Used Kaldi Speech Recognition System, with a public language model, in order to create lattices of thousands of sentences to be used in the interface (which was fed with text-to-speech voices of short 7-word sentences), and Python to conduct processing of sentences to be used
  • Designing a Wizard of Oz study to assess whether word error rate and density affect the amount of effort required to correct speech recognition errors using a graphical representation of a confusion network
[Picture]
Graphical Speech Autocorrection Interface (from Murad et. al, MobileHCI '19 - based on 'Parakeet" by Vertanen et. al, IUI '09). Shade represents confidence level; darker shade = higher confidence. Top row is 1-best result from Confusion Network.

Resulting Publications:

  1. Christine Murad, Cosmin Munteanu, and Wolfgang Stuerzlinger. 2019. Effects of WER on ASR Correction Interfaces for Mobile Text Entry. In Proceedings of the 21st International Conference on Human-Computer Interaction with Mobile Devices and Services (MobileHCI '19). Association for Computing Machinery, New York, NY, USA, Article 56, 1–6. https://doi.org/10.1145/3338286.3344404