copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

Advances in Very Deep Convolutional Neural Networks for LVCSR

T. Sercu, and V. Goel. (Jun 25, 2016)
DOI: 10.21437/Interspeech.2016-1033

Abstract

Very deep CNNs with small 3x3 kernels have recently been shown to achieve very strong performance as acoustic models in hybrid NN-HMM speech recognition systems. In this paper we investigate how to efficiently scale these models to larger datasets. Specifically, we address the design choice of pooling and padding along the time dimension which renders convolutional evaluation of sequences highly inefficient. We propose a new CNN design without timepadding and without timepooling, which is slightly suboptimal for accuracy, but has two significant advantages: it enables sequence training and deployment by allowing efficient convolutional evaluation of full utterances, and, it allows for batch normalization to be straightforwardly adopted to CNNs on sequence data. Through batch normalization, we recover the lost peformance from removing the time-pooling, while keeping the benefit of efficient convolutional evaluation. We demonstrate the performance of our models both on larger scale data than before, and after sequence training. Our very deep CNN model sequence trained on the 2000h switchboard dataset obtains 9.4 word error rate on the Hub5 test-set, matching with a single model the performance of the 2015 IBM system combination, which was the previous best published result.

Links and resources

BibTeX key: sercu2016advances
entry type: misc
year: 2016
month: jun
day: 25
citeulike-article-id: 14172189
eprint: 1604.01792
citeulike-linkout-0: http://dx.doi.org/10.21437/Interspeech.2016-1033
citeulike-linkout-2: http://arxiv.org/pdf/1604.01792
citeulike-linkout-1: http://arxiv.org/abs/1604.01792
archiveprefix: arXiv
priority: 5
posted-at: 2016-10-26 15:39:57
DOI: 10.21437/Interspeech.2016-1033
url: http://dx.doi.org/10.21437/Interspeech.2016-1033

BibSonomy

copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

Advances in Very Deep Convolutional Neural Networks for LVCSR

Abstract

Links and resources

Tags

community

Cite this publication

More citation styles

search on

Meta data

Comments and Reviews
(0)

BibSonomy

copydeleteadd this publication to your clipboardcommunity posthistory of this postURLDOIBibTeXEndNoteAPAChicagoDIN 1505HarvardMSOffice XML Advances in Very Deep Convolutional Neural Networks for LVCSR

Abstract

Links and resources

Tags

community

Cite this publication

More citation styles

search on

Meta data

Comments and Reviews (0)

copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

Advances in Very Deep Convolutional Neural Networks for LVCSR

Comments and Reviews
(0)