Fetching the paper…
Reading the bibliography…
Speech-based inputs have been gaining significant momentum with the popularity of smartphones and tablets in our daily lives, since voice is the most easiest and efficient way for human-computer interaction.
Phoneme recognition using time-delay neural networks
Alex Waibel, Toshiyuki Hanazawa, Geoffrey Hinton, Kiyohiro Shikano, and Kevin J Lang. 1989 · 1989
Earlier work this paper cites.
Near-perfect-reconstruction pseudo-QMF banks
Truong Q Nguyen. 1994 · 1994
Earlier work this paper cites.
Content-based retrieval of music and audio. In Multimedia Storage and Archiving Systems II , Vol. 3229. International Society for Optics and Photonics, 138–147
Jonathan T Foote. 1997 · 1997
Earlier work this paper cites.
Speech reconstruction from mel frequency cepstral coefficients and pitch frequency. In 2000 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No. 00CH37100) , Vol. 3. IEEE, 1299–1302
Dan Chazan, Ron Hoory, Gilad Cohen, and Meir Zibulski. 2000 · 2000
Earlier work this paper cites.
Recurrent neural networks
Larry R Medsker and LC Jain. 2001 · 2001
Earlier work this paper cites.
Demonstration of SpeakQL: Speech-driven Multimodal Querying of Structured Data. In Proceedings of the 2019 International Conference on Management of Data . 2001–2004
Vraj Shah, Side Li, Kevin Yang, Arun Kumar, and Lawrence Saul. 2019 · 2004
Earlier work this paper cites.
Voicecode: An innovative speech interface for programming-by-voice. In CHI’06 extended abstracts on Human factors in computing systems . 239–242
Alain Désilets, David C Fox, and Stuart Norton. 2006 · 2006
Earlier work this paper cites.
TableQA: a Large-Scale Chinese Text-to-SQL Dataset for Table-Aware SQL Generation
Ningyuan Sun, Xuefeng Yang, and Yunfeng Liu. 2020 · 2006
Earlier work this paper cites.
Database Management System
Seema Kedar. 2009 · 2009
Earlier work this paper cites.
From keywords to semantic queries—Incremental query construction on the Semantic Web
Gideon Zenz, Xuan Zhou, Enrico Minack, Wolf Siberski, and Wolfgang Nejdl. 2009 · 2009
Earlier work this paper cites.
Data-thirsty business analysts need SODA: search over data warehouse. In Proceedings of the 20th ACM international conference on Information and knowledge management . 2525–2528
Lukas Blunschi, Claudio Jossen, Donald Kossmann, Magdalini Mori, and Kurt Stockinger. 2011 · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks. In Proceedings of the fourteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 315–323
Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011 · 2011
Earlier work this paper cites.
Natural language interface for database: a brief review
Neelu Nihalani, Sanjay Silakari, and Mahesh Motwani. 2011 · 2011
Earlier work this paper cites.
The Kaldi speech recognition toolkit. In IEEE 2011 workshop on automatic speech recognition and understanding . IEEE Signal Processing Society
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
Verbmobil: foundations of speech-to-speech translation
Wolfgang Wahlster. 2013 · 2013
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. In EMNLP
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
TTS synthesis with bidirectional LSTM based recurrent neural networks. In Fifteenth annual conference of the international speech communication association
Yuchen Fan, Yao Qian, Feng-Long Xie, and Frank K Soong. 2014 · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Constructing an interactive natural language interface for relational databases
Fei Li and HV Jagadish. 2014a · 2014
Earlier work this paper cites.
NaLIR: an interactive natural language interface for querying relational databases. In Proceedings of the 2014 ACM SIGMOD international conference on Management of data . 709–712
Fei Li and Hosagrahar V Jagadish. 2014b · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning . PMLR, 448–456
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation. In EMNLP
Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Sina: Semantic interpretation of user queries for question answering on interlinked data
Saeedeh Shekarpour, Edgard Marx, Axel-Cyrille Ngonga Ngomo, and Sören Auer. 2015 · 2015
Earlier work this paper cites.
Pointer Networks
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
End-to-end attention-based large vocabulary speech recognition. In 2016 ICASSP . IEEE, 4945–4949
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition. In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4960–4964
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. 2016 · 2016
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling. 2016 · 2016
Cited alongside, same era.
Making the case for query-by-voice with echoquery. In Proceedings of the 2016 International Conference on Management of Data . 2129–2132
Gabriel Lyons, Vinh Tran, Carsten Binnig, Ugur Cetintemel, and Tim Kraska. 2016 · 2016
Cited alongside, same era.
Incorporating source syntax into transformer-based neural machine translation. In Proceedings of the Fourth Conference on Machine Translation (Volume 1: Research Papers) . 24–33
Anna Currey and Kenneth Heafield. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Towards Complex Text-to-SQL in Cross-Domain Database with Intermediate Representation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . 4524–4535
Jiaqi Guo, Zecheng Zhan, Yan Gao, Yan Xiao, Jian-Guang Lou, Ting Liu, and Dongmei Zhang. 2019 · 2019
Later among the works it cites.
MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brebisson, Yoshua Bengio, and Aaron Courville. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards speech-to-text translation without speech recognition. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers . 474–479
Sameer Bansal, Herman Kamper, Adam Lopez, and Sharon Goldwater. 2017 · 2017
Cited alongside, same era.
AMUSE: multilingual semantic parsing for question answering over linked data. In International Semantic Web Conference . Springer, 329–346
Sherzod Hakimov, Soufian Jebbara, and Philipp Cimiano. 2017 · 2017
Cited alongside, same era.
Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer. In 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 193–199
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar. 2017 · 2017
Cited alongside, same era.
Voice-based data exploration: Chatting with your database. In Proceedings of the Workshop on Search-Oriented Conversational AI (SCAI)
Prasetya Utama, Nathaniel Weir, Carsten Binnig, and Ugur Cetintemel. 2017 · 2017
Cited alongside, same era.
Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems . 6000–6010
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Hybrid CTC/attention architecture for end-to-end speech recognition
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi. 2017 · 2017
Cited alongside, same era.
Sqlnet: Generating structured queries from natural language without reinforcement learning
Xiaojun Xu, Chang Liu, and Dawn Song. 2017a · 2017
Cited alongside, same era.
Sqlnet: Generating structured queries from natural language without reinforcement learning
Xiaojun Xu, Chang Liu, and Dawn Song. 2017b · 2017
Cited alongside, same era.
Multimodal transformer networks for end-to-end video-grounded dialogue systems.(2019). In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 2019 July 28-August , Vol. 2. 5612–5623
Hung LE, Doyen SAHOO, and Nancy F CHEN. [n.d.] · 2019
Later among the works it cites.
Disability Assistive Programming: Using Voice Input to Write Code
Hunter Lee, James B Fenwick Jr, Richard E Klima, Alice A McRae, and Jefford Vahlbusch. 2019 · 2019
Later among the works it cites.
Deepgcns: Can gcns go as deep as cnns?. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 9267–9276
Guohao Li, Matthias Muller, Ali Thabet, and Bernard Ghanem. 2019 · 2019
Later among the works it cites.
Self-Attention Transducers for End-to-End Speech Recognition
Zhengkun Tian, Jiangyan Yi, Jianhua Tao, Ye Bai, and Zhengqi Wen. 2019 · 2019
Later among the works it cites.
A comparison of Transformer and LSTM encoder decoder models for ASR. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 8–15
Albert Zeyer, Parnia Bahar, Kazuki Irie, Ralf Schlüter, and Hermann Ney. 2019 · 2019
Later among the works it cites.
Editing-Based SQL Query Generation for Cross-Domain Context-Dependent Questions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . 5338–5349
Rui Zhang, Tao Yu, Heyang Er, Sungrok Shim, Eric Xue, Xi Victoria Lin, Tianze Shi, Caiming Xiong, Richard Socher, and Dragomir Radev. 2019 · 2019
Later among the works it cites.
Voxento: A Prototype Voice-controlled Interactive Search Engine for Lifelogs. In Proceedings of the Third Annual Workshop on Lifelog Search Challenge . 77–81
Ahmed Alateeq, Mark Roantree, and Cathal Gurrin. 2020 · 2020
Later among the works it cites.
Handling information loss of graph neural networks for session-based recommendation. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 1172–1180
Tianwen Chen and Raymond Chi-Wing Wong. 2020 · 2020
Later among the works it cites.
Natural language to SQL: Where are we today?
Hyeonji Kim, Byeong-Hoon So, Wook-Shin Han, and Hongrae Lee. 2020 · 2020
Later among the works it cites.
Fastpitch: Parallel text-to-speech with pitch prediction
Adrian Łańcucki. 2020 · 2020
Later among the works it cites.
Re-examining the Role of Schema Linking in Text-to-SQL. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 6943–6954
Wenqiang Lei, Weixin Wang, Zhixin Ma, Tian Gan, Wei Lu, Min-Yen Kan, and Tat-Seng Chua. 2020 · 2020
Later among the works it cites.
Direct speech-to-image translation
Jiguo Li, Xinfeng Zhang, Chuanmin Jia, Jizheng Xu, Li Zhang, Yue Wang, Siwei Ma, and Wen Gao. 2020 · 2020
Later among the works it cites.
Understanding User Perceptions of Robot’s Delay, Voice Quality-Speed Trade-off and GUI during Conversation. In Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems . 1–8
Zhenhui Peng, Kaixiang Mo, Xiaogang Zhu, Junlin Chen, Zhijun Chen, Qian Xu, and Xiaojuan Ma. 2020 · 2020
Later among the works it cites.
Fastspeech 2: Fast and high-quality end-to-end text-to-speech
Yi Ren, Chenxu Hu, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
Athena++ natural language querying for complex nested sql queries
Jaydeep Sen, Chuan Lei, Abdul Quamar, Fatma Özcan, Vasilis Efthymiou, Ayushi Dalmia, Greg Stager, Ashish Mittal, Diptikalyan Saha, and Karthik Sankaranarayanan. 2020 · 2020
Later among the works it cites.
SpeakQL: Towards Speech-driven Multimodal Querying of Structured Data. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data . 2363–2374
Vraj Shah, Side Li, Arun Kumar, and Lawrence Saul. 2020 · 2020
Later among the works it cites.
GoldenRetriever: A Speech Recognition System Powered by Modern Information Retrieval. In MM
Yuanfeng Song, Di Jiang, Xiaoling Huang, Yawen Li, Qian Xu, Raymond Chi Wing Wong, and Qiang Yang. 2020 · 2020
Later among the works it cites.
Demonstrating the voice-based exploration of large data sets with CiceroDB-zero
Immanuel Trummer. 2020 · 2020
Later among the works it cites.
S2IGAN: Speech-to-Image Generation via Adversarial Learning
X Wang, T Qiao, Jihua Zhu, A Hanjalic, and OE Scharenborg. 2020a · 2020
Later among the works it cites.
Multiple knowledge syncretic transformer for natural dialogue generation. In Proceedings of The Web Conference 2020 . 752–762
Xiangyu Zhao, Longbiao Wang, Ruifang He, Ting Yang, Jinxin Chang, and Ruifang Wang. 2020 · 2020
Later among the works it cites.
L2RS: a learning-to-rescore mechanism for automatic speech recognition. In MM
Yuanfeng Song, Di Jiang, Xuefang Zhao, Qian Xu, Raymond Chi-Wing Wong, Lixin Fan, and Qiang Yang. 2021 · 2021
Later among the works it cites.
Generating Images From Spoken Descriptions
Xinsheng Wang, Tingting Qiao, Jihua Zhu, Alan Hanjalic, and Odette Scharenborg. 2021 · 2021
Later among the works it cites.