Fetching the paper…
Reading the bibliography…
Models based on BERT have been extremely successful in solving a variety of natural language processing (NLP) tasks.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books, 2015
Yukun Zhu, Ryan Kiros, Richard Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Deep contextualized word representations, 2018
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Fixup initialization: Residual learning without normalization, 2019
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma · 2019
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention, 2019
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?, 2019
Paul Michel, Omer Levy, and Graham Neubig · 2019
Cited alongside, same era.
Bert rediscovers the classical nlp pipeline, 2019
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Cited alongside, same era.
Understanding and improving layer normalization
Jingjing Xu, Xu Sun, Zhiyuan Zhang, Guangxiang Zhao, and Junyang Lin · 2019
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding, 2019
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Cited alongside, same era.
Patient knowledge distillation for bert model compression, 2019
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu · 2019
Cited alongside, same era.
Language models are few-shot learners, 2020
Albert: A lite bert for self-supervised learning of language representations, 2020
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Later among the works it cites.
Improving transformer models by reordering their sublayers, 2020
Ofir Press, Noah A. Smith, and Omer Levy · 2020
Later among the works it cites.
Improving transformer optimization through better initialization
Xiao Shi Huang, Felipe Perez, Jimmy Ba, and Maksims Volkovs · 2020
Later among the works it cites.
Gaussian error linear units (gelus), 2020
Dan Hendrycks and Kevin Gimpel · 2020
Later among the works it cites.
Bert-of-theseus: Compressing bert by progressive module replacing, 2020
Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou · 2020
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Normalized attention without probability cage, 2020
Oliver Richter and Roger Wattenhofer · 2020
Cited alongside, same era.
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2020
Later among the works it cites.
Softermax: Hardware/software co-design of an efficient softmax for transformers, 2021
Jacob R. Stevens, Rangharajan Venkatesan, Steve Dai, Brucek Khailany, and Anand Raghunathan · 2021
Later among the works it cites.