Fetching the paper…
Reading the bibliography…
Language model pre-training has shown promising results in various downstream tasks.
“Curriculum learning,”
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, · 2009
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Vqa: Visual question answering,”
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh, · 2015
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
G. Hinton, O. Vinyals, and J. Dean, · 2015
Earlier work this paper cites.
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al., · 2016
Earlier work this paper cites.
“Montreal forced aligner: Trainable text-speech alignment using kaldi.,”
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, · 2017
Earlier work this paper cites.
“Towards end-to-end spoken language understanding,”
D. Serdyuk, Y. Wang, C. Fuegen, A. Kumar, B. Liu, and Y. Bengio, · 2018
Earlier work this paper cites.
“Spoken language understanding without speech recognition,”
Y. Chen, R. Price, and S. Bangalore, · 2018
Earlier work this paper cites.
“Speaker recognition from raw waveform with sincnet,”
M. Ravanelli and Y. Bengio, · 2018
Earlier work this paper cites.
“Unsupervised cross-modal alignment of speech and text embedding spaces,”
Y.-A. Chung, W.-H. Weng, S. Tong, and J. Glass, · 2018
Earlier work this paper cites.
A. Coucke, A. Saade, A. Ball, T. Bluche, A. Caulier, D. Leroy, C. Doumouro, T. Gisselbrecht, F. Caltagirone, T. Lavril, et al., · 2018
Cited alongside, same era.
“Spoken language understanding on the edge,”
A. Saade, A. Coucke, A. Caulier, J. Dureau, A. Ball, T. Bluche, D. Leroy, C. Doumouro, T. Gisselbrecht, F. Caltagirone, et al., · 2018
Cited alongside, same era.
“Nsml: Meet the mlaas platform with a real-world case study,”
H. Kim, M. Kim, D. Seo, J. Kim, H. Park, S. Park, H. Jo, K. Kim, Y. Yang, Y. Kim, et al., · 2018
Cited alongside, same era.
“Speech model pre-training for end-to-end spoken language understanding,”
L. Lugosch, M. Ravanelli, P. Ignoto, V. S. Tomar, and Y. Bengio, · 2019
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
“Speechbert: Cross-modal pre-trained language model for end-to-end spoken question answering,”
Y.-S. Chuang, C.-L. Liu, and H.-Y. Lee, · 2019
Later among the works it cites.
“Unicoder: A universal language encoder by pre-training with multiple cross-lingual tasks,”
H. Huang, Y. Liang, N. Duan, M. Gong, L. Shou, D. Jiang, and M. Zhou, · 2019
Later among the works it cites.
“Large-scale unsupervised pre-training for end-to-end spoken language understanding,”
P. Wang, L. Wei, Y. Cao, J. Xie, and Z. Nie, · 2020
Closest in time.
“Speech to text adaptation: Towards an efficient cross-modal distillation,”
W. I. Cho, D. Kwak, J. Yoon, and N. S. Kim, · 2020
Closest in time.
“Pretrained semantic speech embeddings for end-to-end spoken language understanding via cross-modal teacher-student learning,”
P. Denisov and N. T. Vu, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, · 2019
Cited alongside, same era.
“Cross-lingual language model pretraining,”
G. Lample and A. Conneau, · 2019
Cited alongside, same era.
“Recent advances in end-to-end spoken language understanding,”
N. Tomashenko, A. Caubrière, Y. Estève, A. Laurent, and E. Morin, · 2019
Cited alongside, same era.
“End-to-end spoken language understanding: Bootstrapping in low resource scenarios.,”
S. Bhosale, I. Sheikh, S. H. Dumpala, and S. K. Kopparapu, · 2019
Cited alongside, same era.
“Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,”
J. Lu, D. Batra, D. Parikh, and S. Lee, · 2019
Cited alongside, same era.
“From recognition to cognition: Visual commonsense reasoning,”
R. Zellers, Y. Bisk, A. Farhadi, and Y. Choi, · 2019
Cited alongside, same era.
“Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training.,”
G. Li, N. Duan, Y. Fang, M. Gong, D. Jiang, and M. Zhou, · 2020
Closest in time.
“Don’t stop pretraining: Adapt language models to domains and tasks,”
S. Gururangan, A. Marasović, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith, · 2020
Closest in time.
“Curriculum pre-training for end-to-end speech translation,”
C. Wang, Y. Wu, S. Liu, M. Zhou, and Z. Yang, · 2020
Closest in time.
“Learning asr-robust contextualized embeddings for spoken language understanding,”
C.-W. Huang and Y.-N. Chen, · 2020
Closest in time.
“End-to-end spoken language understanding without matched language speech model pretraining data,”
R. Price, · 2020
Closest in time.