Fetching the paper…
Reading the bibliography…
Pre-trained models like BERT (Devlin et al., 2018) have dominated NLP / IR applications such as single sentence classification, text pair classification, and question answering.
Distilling task-specific knowledge from BERT into simple neural networks
R. Tang, Y. Lu, L. Liu, L. Mou, O. Vechtomova, and J. Lin. 2019 · 1903
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O.Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov. 2019 · 1907
Earlier work this paper cites.
To tune or not to tune? how about the best of both worlds?
Ran Wang, Haibo Su, Chunye Wang, Kailin Ji, and Jupeng Ding. 2019 · 1907
Earlier work this paper cites.
Multilingual universal sentence encoder for semantic retrieval
Y. Yang, D. Cer, A. Ahmad, M. Guo, J. Law, N. Constant, G. Hernández Ábrego, S. Yuan, C. Tar, Y. Sung, B. Strope, and R. Kurzweil. 2019 · 1907
Earlier work this paper cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
I. Turc, M. Chang, K. Lee, and K. Toutanova. 2019 · 1908
Earlier work this paper cites.
Tinybert: Distilling BERT for natural language understanding
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu. 2019 · 1909
Earlier work this paper cites.
ALBERT: A lite BERT for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut. 2019 · 1909
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf. 2019 · 1910
Earlier work this paper cites.
An empirical investigation of statistical significance in NLP
T. Berg-Kirkpatrick, D. Burkett, and D. Klein. 2012 · 2012
Earlier work this paper cites.
Learning deep structured semantic models for web search using clickthrough data
P. Huang, X. He, J. Gao, L. Deng, A. Acero, and L. P. Heck. 2013 · 2013
Earlier work this paper cites.
Convolutional neural network architectures for matching natural language sentences
B. Hu, Z. Lu, H. Li, and Q. Chen. 2014 · 2014
Earlier work this paper cites.
Semantic Matching in Search
H. Li and J. Xu. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
S. R. Bowman, G. Angeli, C. Potts, and C. D. Manning. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. E. Hinton, O. Vinyals, and J. Dean. 2015 · 2015
Cited alongside, same era.
Together we stand: Siamese networks for similar question retrieval
A. Das, H. Yenala, M. Chinnakotla, and M. Shrivastava. 2016 · 2016
Cited alongside, same era.
A deep relevance matching model for ad-hoc retrieval
J. Guo, Y. Fan, Q. Ai, and W. B. Croft. 2016 · 2016
Cited alongside, same era.
S. Han, H. Mao, and W. J. Dally. 2016 · 2016
Cited alongside, same era.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and <1mb model size
F. N. Iandola, M. W. Moskewicz, K. Ashraf, S. Han, W.J. Dally, and K. Keutzer. 2016 · 2016
Universal sentence encoder for English
D. Cer, Y. Yang, S. Kong, N. Hua, N. Limtiaco, R. St. John, N. Constant, M. Guajardo-Cespedes, S. Yuan, C. Tar, B. Strope, and R. Kurzweil. 2018 · 2018
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova. 2018 · 2018
Later among the works it cites.
Learning cross-lingual sentence representations via a multi-task dual-encoder model
M. Chidambaram, Y. Yang, D. Cer, S. Yuan, Y. Sung, B. Strope, and R. Kurzweil. 2019 · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin. 2019 · 2019
Later among the works it cites.
A deep look into neural ranking models for information retrieval
J. Guo, Y. Fan, L. Pang, L. Yang, Q. Ai, H. Zamani, C. Wu, W. B. Croft, and X. Cheng. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Text matching as image recognition
L. Pang, Y. Lan, J. Guo, J. Xu, S. Wan, and X. Cheng. 2016 · 2016
Cited alongside, same era.
anmm: Ranking short answer texts with attention-based neural matching model
L. Yang, Q. Ai, J. Guo, and W. B. Croft. 2016 · 2016
Cited alongside, same era.
Efficient natural language response suggestion for smart reply
M. L. Henderson, R. Al-Rfou, B. Strope, Y. Sung, L. Lukács, R. Guo, S. Kumar, B. Miklos, and R. Kurzweil. 2017 · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T.Weyand, M. Andreetto, and H. Adam. 2017 · 2017
Cited alongside, same era.
Learning to match using local and distributed representations of text for web search
B. Mitra, F. Diaz, and N. Craswell. 2017 · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł . Kaiser, and I. Polosukhin. 2017 · 2017
Cited alongside, same era.
A compare-aggregate model for matching text sequences
Shuohang Wang and Jing Jiang. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Bridging the gap between relevance matching and semantic matching for short text similarity modeling
J. Rao, L. Liu, Y. Tay, W. Yang, P. Shi, and J. Lin. 2019 · 2019
Later among the works it cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
N. Reimers and I. Gurevych. 2019 · 2019
Later among the works it cites.
Patient knowledge distillation for BERT model compression
S. Sun, Y. Cheng, Z. Gan, and J. Liu. 2019 · 2019
Later among the works it cites.
ELECTRA: pre-training text encoders as discriminators rather than generators
K. Clark, M. Luong, Q. V. Le, and C. D. Manning. 2020 · 2020
Closest in time.
Poly-encoders: Architectures and pre-training strategies for fast and accurate multi-sentence scoring
S. Humeau, K. Shuster, M. Lachaux, and J. Weston. 2020 · 2020
Closest in time.
Efficient document re-ranking for transformers by precomputing term representations
S. MacAvaney, F. Maria Nardini, R. Perego, N. Tonellotto, N. Goharian, and O. Frieder. 2020 · 2020
Closest in time.
Comparing rewinding and fine-tuning in neural network pruning
A. Renda, J. Frankle, and M.Carbin. 2020 · 2020
Closest in time.