Fetching the paper…
Reading the bibliography…
Knowledge distillation is an effective way to transfer knowledge from a strong teacher to an efficient student model.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Curriculum Learning for Dense Retrieval Distillation. In SIGIR . 1979–1983
Hansi Zeng, Hamed Zamani, and Vishwa Vinay. 2022 · 1983
Earlier work this paper cites.
Overview of the TREC 2019 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2020 · 2003
Earlier work this paper cites.
Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation
Sebastian Hofstätter, Sophia Althammer, Michael Schröder, Mete Sertkan, and Allan Hanbury. 2020 · 2010
Earlier work this paper cites.
Distilling Dense Representations for Ranking using Tightly-Coupled Teachers
Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2020 · 2010
Earlier work this paper cites.
Is Retriever Merely an Approximator of Reader?
Sohee Yang and Minjoon Seo. 2020 · 2010
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
FitNets: Hints for Thin Deep Nets. In ICLR (Poster)
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. In Proceedings of the Workshop on Cognitive Computation: Integrating neural and symbolic approaches 2016 (CEUR Workshop Proceedings, Vol. 1773) , Tarek Richard Besold, Antoine Bordes, Artur S. d’Avila Garcez, and Greg Wayne (Eds.)
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
Efficient Knowledge Distillation from an Ensemble of Teachers. In INTERSPEECH . ISCA, 3697–3701
Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Samuel Thomas, Jia Cui, and Bhuvana Ramabhadran. 2017 · 2017
Earlier work this paper cites.
Learning without Forgetting
Zhizhong Li and Derek Hoiem. 2018 · 2017
Earlier work this paper cites.
Anserini: Enabling the Use of Lucene for Information Retrieval Research. In SIGIR . 1253–1256
Peilin Yang, Hui Fang, and Jimmy Lin. 2017 · 2017
Earlier work this paper cites.
Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer. In ICLR (Poster)
Sergey Zagoruyko and Nikos Komodakis. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Deeper Text Understanding for IR with Contextual Neural Language Modeling. In SIGIR . 985–988
Zhuyun Dai and Jamie Callan. 2019 · 2019
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 2019
Earlier work this paper cites.
Billion-Scale Similarity Search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2021 · 2019
Earlier work this paper cites.
Natural Questions: a Benchmark for Question Answering Research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization. In Proceedings of the 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
Document Expansion by Query Prediction
Rodrigo Frassetto Nogueira, Wei Yang, Jimmy Lin, and Kyunghyun Cho. 2019 · 2019
Cited alongside, same era.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Cited alongside, same era.
Incremental Event Detection via Knowledge Consolidation Networks. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). 707–717
Pengfei Cao, Yubo Chen, Jun Zhao, and Taifeng Wang. 2020 · 2020
Cited alongside, same era.
RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). 2825–2835
Ruiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Qiaoqiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021b · 2021
Later among the works it cites.
Pro-KD: Progressive Distillation by Following the Footsteps of the Teacher
Mehdi Rezagholizadeh, Aref Jafari, Puneeth Salad, Pranav Sharma, Ali Saheb Pasand, and Ali Ghodsi. 2021 · 2021
Later among the works it cites.
End-to-End Training of Neural Retrievers for Open-Domain Question Answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021 , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). Association for Computational Linguistics, 6648–6662
Devendra Singh Sachan, Mostofa Patwary, Mohammad Shoeybi, Neel Kant, Wei Ping, William L. Hamilton, and Bryan Catanzaro. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). 6769–6781
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Cited alongside, same era.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval , Jimmy X. Huang, Yi Chang, Xueqi Cheng, Jaap Kamps, Vanessa Murdock, Ji-Rong Wen, and Yiqun Liu (Eds.). 39–48
Omar Khattab and Matei Zaharia. 2020 · 2020
Cited alongside, same era.
Improved Knowledge Distillation via Teacher Assistant. In AAAI . 5191–5198
Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh. 2020 · 2020
Cited alongside, same era.
Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training
Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
Dataset Cartography: Mapping and Diagnosing Datasets with Training Dynamics. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). 9275–9293
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Understanding Knowledge Distillation in Non-autoregressive Machine Translation. In ICLR
Chunting Zhou, Jiatao Gu, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
Distilling Knowledge via Knowledge Review. In CVPR . 5008–5017
Pengguang Chen, Shu Liu, Hengshuang Zhao, and Jiaya Jia. 2021 · 2021
Cited alongside, same era.
SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval
Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stéphane Clinchant. 2021 · 2021
Cited alongside, same era.
ERNIE-Tiny : A Progressive Distillation Framework for Pretrained Transformer Compression
Weiyue Su, Xuyi Chen, Shikun Feng, Jiaxiang Liu, Weixin Liu, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2021 · 2021
Later among the works it cites.
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. In Proceedings of the 9th International Conference on Learning Representations
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021 · 2021
Later among the works it cites.
Optimizing Dense Retrieval Model Training with Hard Negatives. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , Fernando Diaz, Chirag Shah, Torsten Suel, Pablo Castells, Rosie Jones, and Tetsuya Sakai (Eds.). 1503–1512
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. 2021 · 2021
Later among the works it cites.
SPARTA: Efficient Open-Domain Question Answering via Sparse Transformer Matching Retrieval. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tür, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou (Eds.). 565–575
Tiancheng Zhao, Xiaopeng Lu, and Kyusong Lee. 2021 · 2021
Later among the works it cites.
Robust Cross-Modal Representation Learning with Progressive Self-Distillation. In CVPR . IEEE, 16409–16420
Alex Andonian, Shixing Chen, and Raffay Hamid. 2022 · 2022
Closest in time.
Understanding Dataset Difficulty with V -Usable Information. In Proceedings of the 39th International Conference on Machine Learning , Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato (Eds.), Vol. 162. 5988–6008
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. 2022 · 2022
Closest in time.
Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune Paradigm. In ACL (1) . Association for Computational Linguistics, 190–200
Shaoyi Huang, Dongkuan Xu, Ian En-Hsu Yen, Yijue Wang, Sung-En Chang, Bingbing Li, Shiyang Chen, Mimi Xie, Sanguthevar Rajasekaran, Hang Liu, and Caiwen Ding. 2022 · 2022
Closest in time.
Yuxiang Lu, Yiding Liu, Jiaxiang Liu, Yunsheng Shi, Zhengjie Huang, Shikun Feng, Yu Sun, Hao Tian, Hua Wu, Shuaiqiang Wang, Dawei Yin, and Haifeng Wang. 2022 · 2022
Closest in time.
Domain-matched Pre-training Tasks for Dense Retrieval. In Findings of the Association for Computational Linguistics: NAACL 2022, Seattle, WA, United States, July 10-15, 2022 , Marine Carpuat, Marie-Catherine de Marneffe, and Iván Vladimir Meza Ruíz (Eds.). Association for Computational Linguistics, 1524–1534
Barlas Oguz, Kushal Lakhotia, Anchit Gupta, Patrick Lewis, Vladimir Karpukhin, Aleksandra Piktus, Xilun Chen, Sebastian Riedel, Scott Yih, Sonal Gupta, and Yashar Mehdad. 2022 · 2022
Closest in time.
Progressive Distillation for Fast Sampling of Diffusion Models. In ICLR . OpenReview.net
Tim Salimans and Jonathan Ho. 2022 · 2022
Closest in time.
ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2022, Seattle, WA, United States, July 10-15, 2022 , Marine Carpuat, Marie-Catherine de Marneffe, and Iván Vladimir Meza Ruíz (Eds.). Association for Computational Linguistics, 3715–3734
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022 · 2022
Closest in time.
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al · 2022
Closest in time.
Unified and Effective Ensemble Knowledge Distillation
Chuhan Wu, Fangzhao Wu, Tao Qi, and Yongfeng Huang. 2022 · 2022
Closest in time.
LED: Lexicon-Enlightened Dense Retriever for Large-Scale Retrieval
Kai Zhang, Chongyang Tao, Tao Shen, Can Xu, Xiubo Geng, Binxing Jiao, and Daxin Jiang. 2022b · 2022
Closest in time.
BERT Learns to Teach: Knowledge Distillation with Meta Learning. In ACL (1) . 7037–7049
Wangchunshu Zhou, Canwen Xu, and Julian J. McAuley. 2022 · 2022
Closest in time.