Autocorrection of Speech Recognition Errors Using Confusion Network GUI
University of Toronto (Jan 2017 - Jan 2020)
Project Summary
Developed a graphical web interface in Node.JS that shows a visual representation of a speech recognition confusion network, from a lattice, that can be manipulated through touch in order to correct speech recognition errors (based on “Parakeet” by Vertanen et. al [IUI ‘09])
Used Kaldi Speech Recognition System, with a public language model, in order to create lattices of thousands of sentences to be used in the interface (which was fed with text-to-speech voices of short 7-word sentences), and Python to conduct processing of sentences to be used
Designing a Wizard of Oz study to assess whether word error rate and density affect the amount of effort required to correct speech recognition errors using a graphical representation of a confusion network
[Picture]
Graphical Speech Autocorrection Interface (from Murad et. al, MobileHCI '19 - based on 'Parakeet" by Vertanen et. al, IUI '09). Shade represents confidence level; darker shade = higher confidence. Top row is 1-best result from Confusion Network.
Resulting Publications:
Christine Murad, Cosmin Munteanu, and Wolfgang Stuerzlinger. 2019. Effects of WER on ASR Correction Interfaces for Mobile Text Entry. In Proceedings of the 21st International Conference on Human-Computer Interaction with Mobile Devices and Services (MobileHCI '19). Association for Computing Machinery, New York, NY, USA, Article 56, 1–6. https://doi.org/10.1145/3338286.3344404