Fetching the paper…
Reading the bibliography…
Causal Language Modeling (CLM) and Masked Language Modeling (MLM) are two mainstream learning paradigms based on Transformer networks, specifically the Decoder-only and Encoder-only architectures.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu. 2019 · 1907
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Z Lan. 2019 · 1909
Earlier work this paper cites.
“cloze procedure”: A new tool for measuring readability
Wilson L Taylor. 1953 · 1953
Earlier work this paper cites.
n-gram statistics for natural language understanding and text processing
Ching Y. Suen. 1979 · 1979
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman. 1990 · 1990
Earlier work this paper cites.
Long short-term memory
S Hochreiter. 1997 · 1997
Earlier work this paper cites.
The CHILDES project: The database , volume 2
Brian MacWhinney. 2000 · 2000
Earlier work this paper cites.
Glu variants improve transformer
Noam Shazeer. 2020 · 2002
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020 · 2006
Earlier work this paper cites.
The amara corpus: Building parallel language resources for the educational domain
Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, and Stephan Vogel. 2014 · 2014
Earlier work this paper cites.
The Goldilocks principle: Reading children’s books with explicit memory representations
Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. 2016 · 2016
Cited alongside, same era.
Opensubtitles2016: Extracting large parallel corpora from movie and tv subtitles
Pierre Lison and Jörg Tiedemann. 2016 · 2016
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford. 2018 · 2018
Cited alongside, same era.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 2019
Cited alongside, same era.
Normformer: Improved transformer pretraining with extra normalization
Sam Shleifer, Jason Weston, and Myle Ott. 2021 · 2021
Later among the works it cites.
Not all layers are equally as important: Every layer counts bert
Lucas Georges Gabriel Charpentier and David Samuel. 2023 · 2023
Later among the works it cites.
How good are gpt models at machine translation? a comprehensive evaluation
Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023 · 2023
Later among the works it cites.
End-to-end speech recognition: A survey
Rohit Prabhavalkar, Takaaki Hori, Tara N Sainath, Ralf Schlüter, and Shinji Watanabe. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
A hierarchical multi-task approach for learning embeddings from semantic tasks
Victor Sanh, Thomas Wolf, and Sebastian Ruder. 2019 · 2019
Cited alongside, same era.
A standardized project gutenberg corpus for statistical analysis of natural language and quantitative linguistics
Martin Gerlach and Francesc Font-Clos. 2020 · 2020
Cited alongside, same era.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Cited alongside, same era.
Attention-informed mixed-language training for zero-shot cross-lingual task-oriented dialogue systems
Zihan Liu, Genta Indra Winata, Zhaojiang Lin, Peng Xu, and Pascale Fung. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Alex Warstadt, Leshem Choshen, Aaron Mueller, Adina Williams, Ethan Wilcox, and Chengxu Zhuang. 2023a
Cited in the paper.
David Samuel. 2023 · 2023
Later among the works it cites.
Trained on 100 million words and still in shape: Bert meets british national corpus
David Samuel, Andrey Kutuzov, Lilja Øvrelid, and Erik Velldal. 2023 · 2023
Later among the works it cites.
A survey of sentiment analysis: Approaches, datasets, and future research
Kian Long Tan, Chin Poo Lee, and Kian Ming Lim. 2023 · 2023
Later among the works it cites.
Inar Timiryasov and Jean-Loup Tastet. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Later among the works it cites.
Zero-shot information extraction via chatting with chatgpt
Xiang Wei, Xingyu Cui, Ning Cheng, Xiaobin Wang, Xin Zhang, Shen Huang, Pengjun Xie, Jinan Xu, Yufeng Chen, Meishan Zhang, et al. 2023 · 2023
Later among the works it cites.
Leshem Choshen, Ryan Cotterell, Michael Y. Hu, Tal Linzen, Aaron Mueller, Candace Ross, Alex Warstadt, Ethan Wilcox, Adina Williams, and Chengxu Zhuang. 2024 · 2024
Closest in time.