Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon. 2019 · 2019
Later among the works it cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019a · 2019
Later among the works it cites.
Eda: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou. 2019 · 2019
Later among the works it cites.
Conditional bert contextual augmentation
Xing Wu, Shangwen Lv, Liangjun Zang, Jizhong Han, and Songlin Hu. 2019 · 2019
Later among the works it cites.
Be your own teacher: Improve the performance of convolutional neural networks via self distillation
Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma. 2019 · 2019
Later among the works it cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020 · 2020
Later among the works it cites.
Self-training avoids using spurious features under domain shift
Yining Chen, Colin Wei, Ananya Kumar, and Tengyu Ma. 2020b · 2020
Later among the works it cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Later among the works it cites.
Randaugment: Practical automated data augmentation with a reduced search space
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. 2020 · 2020
Later among the works it cites.
A knowledge-enhanced pretraining model for commonsense story generation
Jian Guan, Fei Huang, Zhihao Zhao, Xiaoyan Zhu, and Minlie Huang. 2020 · 2020
Later among the works it cites.
Self-distillation amplifies regularization in hilbert space
Hossein Mobahi, Mehrdad Farajtabar, and Peter L. Bartlett. 2020 · 2020
Later among the works it cites.
Facts2Story: Controlling text generation by key facts
Eyal Orbach and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Later among the works it cites.
Understanding and mitigating the tradeoff between robustness and accuracy
Aditi Raghunathan, Sang Michael Xie, Fanny Yang, John Duchi, and Percy Liang. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Self-training with noisy student improves imagenet classification
Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le. 2020 · 2020
Later among the works it cites.
Bert-of-theseus: Compressing bert by progressive module replacing
Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou. 2020 · 2020
Later among the works it cites.
A survey of data augmentation approaches for NLP
Steven Y. Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard Hovy. 2021 · 2021
Closest in time.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021 · 2021
Closest in time.
Improving robustness using generated data
Sven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg, Dan Andrei Calian, and Timothy A Mann. 2021 · 2021
Closest in time.
Scaling laws for transfer
Original
Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish. 2021 · 2021
Closest in time.
Mate-kd: Masked adversarial text, a companion to knowledge distillation
Original
Ahmad Rashid, Vasileios Lioutas, and Mehdi Rezagholizadeh. 2021 · 2021
Closest in time.
Generalised unsupervised domain adaptation of neural machine translation with cross-lingual data selection
Thuy Vu, Xuanli He, Dinh Phung, and Gholamreza Haffari. 2021a · 2021
Closest in time.
Strata: Self-training with task augmentation for better few-shot learning
Tu Vu, Minh-Thang Luong, Quoc Le, Grady Simon, and Mohit Iyyer. 2021b · 2021
Closest in time.
Learning to synthesize data for semantic parsing
Bailin Wang, Wenpeng Yin, Xi Victoria Lin, and Caiming Xiong. 2021 · 2021
Closest in time.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Closest in time.
Theoretical analysis of self-training with deep networks on unlabeled data
Colin Wei, Kendrick Shen, Yining Chen, and Tengyu Ma. 2021 · 2021
Closest in time.