Fetching the paper…
Reading the bibliography…
BERT, which stands for Bidirectional Encoder Representations from Transformers, is a recently introduced language representation model based upon the transfer learning paradigm.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Classification using discriminative restricted boltzmann machines,”
Hugo Larochelle and Yoshua Bengio, · 2008
Earlier work this paper cites.
“Towards real-time measurement of customer satisfaction using automatically generated call transcripts,”
Youngja Park and Stephen C Gates, · 2009
Earlier work this paper cites.
“Replicated softmax: an undirected topic model,”
Geoffrey E Hinton and Ruslan R Salakhutdinov, · 2009
Earlier work this paper cites.
“Mce training techniques for topic identification of spoken audio documents,”
Timothy J Hazen, · 2011
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Character-level convolutional networks for text classification,”
Xiang Zhang, Junbo Zhao, and Yann LeCun, · 2015
Earlier work this paper cites.
“Learning document representations using subspace multinomial model.,”
Santosh Kesiraju, Lukás Burget, Igor Szöke, and Jan Cernockỳ, · 2016
Cited alongside, same era.
“Predicting user satisfaction from turn-taking in spoken conversations.,”
Shammur Absar Chowdhury, Evgeny A Stepanov, Giuseppe Riccardi, et al., · 2016
Cited alongside, same era.
“Can machine learning techniques predict customer dissatisfaction? a feasibility study for the automotive industry,”
Stefan Meinzer, Ulf Jensen, Alexander Thamm, Joachim Hornegger, and Björn M Eskofier, · 2016
Cited alongside, same era.
“Hierarchical attention networks for document classification,”
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy, · 2016
Cited alongside, same era.
“Purely sequence-trained neural networks for asr based on lattice-free mmi.,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Cited alongside, same era.
“Kate: K-competitive autoencoder for text,”
Yu Chen and Mohammed J Zaki, · 2017
Later among the works it cites.
“Scdv: Sparse composite document vectors using soft clustering over distributional representations,”
Dheeraj Mekala, Vivek Gupta, Bhargavi Paranjape, and Harish Karnick, · 2017
Later among the works it cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Later among the works it cites.
“Joint verification-identification in end-to-end multi-scale cnn framework for topic identification,”
Raghavendra Pappagari, Jesús Villalba, and Najim Dehak, · 2018
Later among the works it cites.
“Long length document classification by local convolutional feature aggregation,”
Liu Liu, Kaile Liu, Zhenghai Cong, Jiali Zhao, Yefei Ji, and Jun He, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“The role of linguistic and prosodic cues on the prediction of self-reported satisfaction in contact centre phone calls,”
Jordi Luque, Carlos Segura, Ariadna Sánchez, Martı Umbert, and Luis Angel Galindo, · 2017
Cited alongside, same era.
Later among the works it cites.
“Transformer-xl: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov, · 2019
Closest in time.
“Docbert: Bert for document classification,”
Ashutosh Adhikari, Achyudh Ram, Raphael Tang, and Jimmy Lin, · 2019
Closest in time.