Understand
Deep language models such as BERT pre-trained on large corpus have given a huge performance boost to the state-of-the-art information retrieval ranking systems.
- Knowledge embedded in such models allows them to pick up complex matching signals between passages and queries.
- However, the high computation cost during inference limits their deployment in real-world search scenarios.
- In this paper, we study if and how the knowledge for search within BERT can be transferred to a smaller ranker through distillation.
Built on
Distilling the Knowledge in a Neural Network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Overview of the TREC 2019 deep learning track. In TREC (to appear)
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2019 · 2019
Earlier work this paper cites.
Similar
Deeper Text Understanding for IR with Contextual Neural Language Modeling. In The 42nd International ACM SIGIR Conference on Research & Development in Information Retrieval
Zhuyun Dai and Jamie Callan. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
TinyBERT: Distilling BERT for Natural Language Understanding
Xiaoqi Jiao, Y. Yin, Lifeng Shang, Xin Jiang, Xusong Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 2019
Cited alongside, same era.
Rodrigo Nogueira and Kyunghyun Cho. 2019 · 2019
Cited alongside, same era.
Then
Document expansion by query prediction
Rodrigo Nogueira, Wei Yang, Jimmy Lin, and Kyunghyun Cho. 2019 · 2019
Later among the works it cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
Patient Knowledge Distillation for BERT Model Compression. In EMNLP/IJCNLP
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Later among the works it cites.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…