Fetching the paper…
Reading the bibliography…
Knowledge distillation is often used to transfer knowledge from a strong teacher model to a relatively weak student model.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Curriculum Learning for Dense Retrieval Distillation. In SIGIR . 1979–1983
Hansi Zeng, Hamed Zamani, and Vishwa Vinay. 2022 · 1983
Earlier work this paper cites.
Overview of the TREC 2019 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2020 · 2003
Earlier work this paper cites.
Curriculum learning. In Proceedings of the 26th annual international conference on machine learning . 41–48
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation
Sebastian Hofstätter, Sophia Althammer, Michael Schröder, Mete Sertkan, and Allan Hanbury. 2020 · 2010
Earlier work this paper cites.
Distilling Dense Representations for Ranking using Tightly-Coupled Teachers
Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2020a · 2010
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
FitNets: Hints for Thin Deep Nets. In ICLR (Poster)
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. In Proceedings of the Workshop on Cognitive Computation: Integrating neural and symbolic approaches 2016 (CEUR Workshop Proceedings, Vol. 1773) , Tarek Richard Besold, Antoine Bordes, Artur S. d’Avila Garcez, and Greg Wayne (Eds.)
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
Like what you like: Knowledge distill via neuron selectivity transfer
Zehao Huang and Naiyan Wang. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Anserini: Enabling the Use of Lucene for Information Retrieval Research. In SIGIR . 1253–1256
Peilin Yang, Hui Fang, and Jimmy Lin. 2017 · 2017
Earlier work this paper cites.
Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer. In ICLR (Poster)
Sergey Zagoruyko and Nikos Komodakis. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Paraphrasing complex network: Network compression via factor transfer
Jangho Kim, SeongUk Park, and Nojun Kwak. 2018 · 2018
Earlier work this paper cites.
Learning deep representations with probabilistic knowledge transfer. In Proceedings of the European Conference on Computer Vision (ECCV) . 268–284
Nikolaos Passalis and Anastasios Tefas. 2018 · 2018
Earlier work this paper cites.
Deeper Text Understanding for IR with Contextual Neural Language Modeling. In SIGIR . 985–988
Zhuyun Dai and Jamie Callan. 2019 · 2019
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 2019
Earlier work this paper cites.
Document Expansion by Query Prediction
Rodrigo Frassetto Nogueira, Wei Yang, Jimmy Lin, and Kyunghyun Cho. 2019 · 2019
Earlier work this paper cites.
Towards understanding knowledge distillation. In International conference on machine learning . PMLR, 5142–5151
Mary Phuong and Christoph Lampert. 2019 · 2019
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Cited alongside, same era.
A survey on image data augmentation for deep learning
Connor Shorten and Taghi M Khoshgoftaar. 2019 · 2019
Cited alongside, same era.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019a · 2019
Cited alongside, same era.
ERNIE 2.0: A Continual Pre-training Framework for Language Understanding
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang. 2019b · 2019
Cited alongside, same era.
Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). 6769–6781
In-batch negatives for knowledge distillation with tightly-coupled teachers for dense retrieval. In Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2021) . 163–173
Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2021 · 2021
Later among the works it cites.
Sparse, dense, and attentional representations for text retrieval
Yi Luan, Jacob Eisenstein, Kristina Toutanova, and Michael Collins. 2021 · 2021
Later among the works it cites.
Unsupervised corpus aware language model pre-training for dense passage retrieval
Gao Luyu and Callan Jamie. 2021 · 2021
Later among the works it cites.
Large dual encoders are generalizable retrievers
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y Zhao, Yi Luan, Keith B Hall, Ming-Wei Chang, et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Cited alongside, same era.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval , Jimmy X. Huang, Yi Chang, Xueqi Cheng, Jaap Kamps, Vanessa Murdock, Ji-Rong Wen, and Yiqun Liu (Eds.). 39–48
Omar Khattab and Matei Zaharia. 2020 · 2020
Cited alongside, same era.
Distilling dense representations for ranking using tightly-coupled teachers
Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2020b · 2020
Cited alongside, same era.
Improved knowledge distillation via teacher assistant. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 5191–5198
Seyed Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh. 2020 · 2020
Cited alongside, same era.
Understanding and improving knowledge distillation
Jiaxi Tang, Rakesh Shivanna, Zhe Zhao, Dong Lin, Anima Singh, Ed H Chi, and Sagar Jain. 2020 · 2020
Cited alongside, same era.
Exclusivity-consistency regularized knowledge distillation for face recognition. In European Conference on Computer Vision . Springer, 325–342
Xiaobo Wang, Tianyu Fu, Shengcai Liao, Shuo Wang, Zhen Lei, and Tao Mei. 2020 · 2020
Cited alongside, same era.
Feature normalized knowledge distillation for image classification. In European Conference on Computer Vision . Springer, 664–680
Kunran Xu, Lai Rui, Yishi Li, and Lin Gu. 2020 · 2020
Cited alongside, same era.
Overview of the TREC 2020 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2021 · 2021
Cited alongside, same era.
Peyman Passban, Yimeng Wu, Mehdi Rezagholizadeh, and Qun Liu. 2021 · 2021
Later among the works it cites.
RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering. In NAACL-HLT . 5835–5847
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021 · 2021
Later among the works it cites.
PAIR: Leveraging Passage-Centric Similarity Relation for Improving Dense Passage Retrieval. In Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021 (Findings of ACL, Vol. ACL/IJCNLP 2021) , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). 2173–2183
Ruiyang Ren, Shangwen Lv, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Qiaoqiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021a · 2021
Later among the works it cites.
RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). 2825–2835
Ruiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Qiaoqiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021b · 2021
Later among the works it cites.
Simple entity-centric questions challenge dense retrievers
Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen. 2021 · 2021
Later among the works it cites.
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. In Proceedings of the 9th International Conference on Learning Representations
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021 · 2021
Later among the works it cites.
Optimizing Dense Retrieval Model Training with Hard Negatives. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , Fernando Diaz, Chirag Shah, Torsten Suel, Pablo Castells, Rosie Jones, and Tetsuya Sakai (Eds.). 1503–1512
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. 2021 · 2021
Later among the works it cites.
Inpars: Data augmentation for information retrieval using large language models
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022 · 2022
Closest in time.
RAIL-KD: RAndom Intermediate Layer Mapping for Knowledge Distillation. In Findings of the Association for Computational Linguistics: NAACL 2022 , Marine Carpuat, Marie-Catherine de Marneffe, and Iván Vladimir Meza Ruíz (Eds.). Association for Computational Linguistics, 1389–1400
Md. Akmal Haidar, Nithin Anchuri, Mehdi Rezagholizadeh, Abbas Ghaddar, Philippe Langlais, and Pascal Poupart. 2022 · 2022
Closest in time.
Curriculum Sampling for Dense Retrieval with Document Expansion
Xingwei He, Yeyun Gong, A Jin, Hang Zhang, Anlei Dong, Jian Jiao, Siu Ming Yiu, Nan Duan, et al · 2022
Closest in time.
PROD: Progressive Distillation for Dense Retrieval
Zhenghao Lin, Yeyun Gong, Xiao Liu, Hang Zhang, Chen Lin, Anlei Dong, Jian Jiao, Jingwen Lu, Daxin Jiang, Rangan Majumder, et al · 2022
Closest in time.
Yuxiang Lu, Yiding Liu, Jiaxiang Liu, Yunsheng Shi, Zhengjie Huang, Shikun Feng, Yu Sun, Hao Tian, Hua Wu, Shuaiqiang Wang, Dawei Yin, and Haifeng Wang. 2022 · 2022
Closest in time.
ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , Marine Carpuat, Marie-Catherine de Marneffe, and Iván Vladimir Meza Ruíz (Eds.). Association for Computational Linguistics, 3715–3734
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022 · 2022
Closest in time.
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al · 2022
Closest in time.
SimANS: Simple Ambiguous Negatives Sampling for Dense Text Retrieval
Kun Zhou, Yeyun Gong, Xiao Liu, Wayne Xin Zhao, Yelong Shen, Anlei Dong, Jingwen Lu, Rangan Majumder, Ji-Rong Wen, Nan Duan, and Weizhu Chen. 2022a · 2022
Closest in time.
MASTER: Multi-task Pre-trained Bottlenecked Masked Autoencoders are Better Dense Retrievers
Kun Zhou, Xiao Liu, Yeyun Gong, Wayne Xin Zhao, Daxin Jiang, Nan Duan, and Ji-Rong Wen. 2022b · 2022
Closest in time.