2020

Understanding BERT Rankers Under Distillation

Gao, Luyu, Dai, Zhuyun, Callan, Jamie

Understand

Deep language models such as BERT pre-trained on large corpus have given a huge performance boost to the state-of-the-art information retrieval ranking systems.

  • Knowledge embedded in such models allows them to pick up complex matching signals between passages and queries.
  • However, the high computation cost during inference limits their deployment in real-world search scenarios.
  • In this paper, we study if and how the knowledge for search within BERT can be transferred to a smaller ranker through distillation.

Built on

Similar

  • Deeper Text Understanding for IR with Contextual Neural Language Modeling. In The 42nd International ACM SIGIR Conference on Research & Development in Information Retrieval

    Zhuyun Dai and Jamie Callan. 2019 · 2019

    Cited alongside, same era.

  • BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019

    Cited alongside, same era.

  • TinyBERT: Distilling BERT for Natural Language Understanding

    Original

    Xiaoqi Jiao, Y. Yin, Lifeng Shang, Xin Jiang, Xusong Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 2019

    Cited alongside, same era.

  • Passage Re-ranking with BERT

    Original

    Rodrigo Nogueira and Kyunghyun Cho. 2019 · 2019

    Cited alongside, same era.

Then

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…