Fetching the paper…
Reading the bibliography…
Knowledge distillation, the technique of transferring knowledge from large, complex models to smaller ones, marks a pivotal step towards efficient AI deployment.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
An information bottleneck approach for controlling conciseness in rationale extraction
Bhargavi Paranjape, Mandar Joshi, John Thickstun, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020 · 1952
Earlier work this paper cites.
Multitask learning
Rich Caruana. 1997 · 1997
Earlier work this paper cites.
Elements of information theory
Thomas M Cover. 1999 · 1999
Earlier work this paper cites.
The information bottleneck: Theory and applications
Noam Slonim. 2002 · 2002
Earlier work this paper cites.
Feature selection with dynamic mutual information
Huawen Liu, Jigui Sun, Lei Liu, and Huijie Zhang. 2009 · 2009
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky. 2015 · 2015
Earlier work this paper cites.
Deep variational information bottleneck
Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy. 2016 · 2016
Earlier work this paper cites.
Mutual information neural estimation
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeshwar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, and Devon Hjelm. 2018 · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla. 2018 · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2018 · 2018
Earlier work this paper cites.
Learning representations by maximizing mutual information across views
Philip Bachman, R Devon Hjelm, and William Buchwalter. 2019 · 2019
Earlier work this paper cites.
Pareto multi-task learning
Xi Lin, Hui-Ling Zhen, Zhenhua Li, Qing-Fu Zhang, and Sam Kwong. 2019 · 2019
Earlier work this paper cites.
Multi-task deep neural networks for natural language understanding
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2019 · 2019
Earlier work this paper cites.
Towards understanding knowledge distillation
Mary Phuong and Christoph Lampert. 2019 · 2019
Earlier work this paper cites.
On variational bounds of mutual information
Ben Poole, Sherjil Ozair, Aaron Van Den Oord, Alex Alemi, and George Tucker. 2019 · 2019
Earlier work this paper cites.
On mutual information maximization for representation learning
Michael Tschannen, Josip Djolonga, Paul K Rubenstein, Sylvain Gelly, and Mario Lucic. 2019 · 2019
Cited alongside, same era.
Deep multi-view information bottleneck
Qi Wang, Claire Boudreau, Qixing Luo, Pang-Ning Tan, and Jiayu Zhou. 2019 · 2019
Cited alongside, same era.
Learning variational word masks to improve the interpretability of neural text classifiers
Hanjie Chen and Yangfeng Ji. 2020 · 2020
Cited alongside, same era.
Learning to learn with variational information bottleneck for domain generalization
Yingjun Du, Jun Xu, Huan Xiong, Qiang Qiu, Xiantong Zhen, Cees GM Snoek, and Ling Shao. 2020 · 2020
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Cited alongside, same era.
Knowledge distillation for multi-task learning
Self-distillation: Towards efficient and compact neural networks
Linfeng Zhang, Chenglong Bao, and Kaisheng Ma. 2021 · 2021
Later among the works it cites.
A survey on multi-task learning
Yu Zhang and Qiang Yang. 2021 · 2021
Later among the works it cites.
Hard gate knowledge distillation-leverage calibration for robust and reliable language model
Dongkyu Lee, Zhiliang Tian, Yingxiu Zhao, Ka Chun Cheung, and Nevin Zhang. 2022 · 2022
Later among the works it cites.
Do text-to-text multi-task learners suffer from task conflict?
David Mueller, Nicholas Andrews, and Mark Dredze. 2022 · 2022
Later among the works it cites.
Cross-task knowledge distillation in multi-task recommendation
Chenxiao Yang, Junwei Pan, Xiaofeng Gao, Tingyu Jiang, Dapeng Liu, and Guihai Chen. 2022 · 2022
Later among the works it cites.
Improving the adversarial robustness of nlp models by information bottleneck
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei-Hong Li and Hakan Bilen. 2020 · 2020
Cited alongside, same era.
Formal limitations on the measurement of mutual information
David McAllester and Karl Stratos. 2020 · 2020
Cited alongside, same era.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020 · 2020
Cited alongside, same era.
Multi-task learning for natural language processing in the 2020s: where are we going?
Joseph Worsham and Jugal Kalita. 2020 · 2020
Cited alongside, same era.
Distilling knowledge via knowledge review
Pengguang Chen, Shu Liu, Hengshuang Zhao, and Jiaya Jia. 2021 · 2021
Cited alongside, same era.
Cenyuan Zhang, Xiang Zhou, Yixin Wan, Xiaoqing Zheng, Kai-Wei Chang, and Cho-Jui Hsieh. 2022a · 2022
Later among the works it cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li. 2023 · 2023
Later among the works it cites.
A close look into the calibration of pre-trained language models
Yangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu, and Heng Ji. 2023 · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023 · 2023
Later among the works it cites.
Learning to maximize mutual information for dynamic feature selection
Ian Connick Covert, Wei Qiu, Mingyu Lu, Na Yoon Kim, Nathan J White, and Su-In Lee. 2023 · 2023
Later among the works it cites.
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Cheng-Yu Hsieh, Chun-Liang Li, Chih-kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023 · 2023
Later among the works it cites.
Gpt-4 as an effective zero-shot evaluator for scientific figure captions
Ting-Yao Hsu, Chieh-Yang Huang, Ryan Rossi, Sungchul Kim, C Giles, and Ting-Hao Huang. 2023 · 2023
Later among the works it cites.
Symbolic chain-of-thought distillation: Small models can also" think" step-by-step
Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, and Yejin Choi. 2023 · 2023
Later among the works it cites.
Gpteval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Later among the works it cites.
Yuhan Ma, Haiqi Jiang, and Chenyou Fan. 2023 · 2023
Later among the works it cites.
Teaching small language models to reason
Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, and Aliaksei Severyn. 2023 · 2023
Later among the works it cites.
Multi-task learning with knowledge distillation for dense prediction
Yangyang Xu, Yibo Yang, and Lefei Zhang. 2023 · 2023
Later among the works it cites.
A survey of multi-task learning in natural language processing: Regarding task relatedness and training methods
Zhihan Zhang, Wenhao Yu, Mengxia Yu, Zhichun Guo, and Meng Jiang. 2023 · 2023
Later among the works it cites.