Fetching the paper…
Reading the bibliography…
We present the submission of the ILLC at the University of Amsterdam to the BabyLM challenge (Warstadt et al., 2023), in the strict-small track.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
The CHILDES project: The database , volume 2
Brian MacWhinney. 2000 · 2000
Earlier work this paper cites.
Dialogue act modeling for automatic tagging and recognition of conversational speech
Andreas Stolcke, Klaus Ries, Noah Coccaro, Elizabeth Shriberg, Rebecca Bates, Daniel Jurafsky, Paul Taylor, Rachel Martin, Carol Van Ess-Dykema, and Marie Meteer. 2000 · 2000
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Natural language from artificial life
Simon Kirby. 2002 · 2002
Earlier work this paper cites.
Transferring Inductive Biases through Knowledge Distillation
Samira Abnar, Mostafa Dehghani, and Willem Zuidema. 2020 · 2006
Earlier work this paper cites.
Emergent multi-agent communication in the deep learning era
Angeliki Lazaridou and Marco Baroni. 2020 · 2006
Earlier work this paper cites.
Nur Ahmed and Muntasir Wahed. 2020 · 2010
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
The amara corpus: Building parallel language resources for the educational domain
Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, and Stephan Vogel. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
The goldilocks principle: Reading children’s books with explicit memory representations
Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. 2015 · 2015
Earlier work this paper cites.
Compression and communication in the cultural evolution of linguistic structure
Simon Kirby, Monica Tamariz, Hannah Cornish, and Kenny Smith. 2015 · 2015
Earlier work this paper cites.
Opensubtitles2016: Extracting large parallel corpora from movie and tv subtitles
Pierre Lison and Jörg Tiedemann. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Emergence of language with multi-agent games: Learning to communicate with sequences of symbols
Serhii Havrylov and Ivan Titov. 2017 · 2017
Earlier work this paper cites.
Natural language does not emerge ‘naturally’in multi-agent dialog
Satwik Kottur, José Moura, Stefan Lee, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Hierarchical representation and estimation of prosody using continuous wavelet transform
Antti Suni, Juraj Šimko, Daniel Aalto, and Martti Vainio. 2017 · 2017
Earlier work this paper cites.
How agents see things: On visual representations in an emergent language game
Diane Bouchacourt and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Martin Gerlach and Francesc Font-Clos. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Unsupervised latent tree induction with deep inside-outside recursive autoencoders
Andrew Drozdov, Pat Verga, Mohit Yadav, Mohit Iyyer, and Andrew McCallum. 2019 · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
An embarrassingly simple method to mitigate undesirable properties of pretrained language model tokenizers
Valentin Hofmann, Hinrich Schuetze, and Janet Pierrehumbert. 2022 · 2022
Later among the works it cites.
Question Answering Infused Pre-training of General-Purpose Contextualized Representations
Robin Jia, Mike Lewis, and Luke Zettlemoyer. 2022 · 2022
Later among the works it cites.
Deduplicating Training Data Makes Language Models Better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022 · 2022
Later among the works it cites.
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas Nolte, Eric J Michaud, Max Tegmark, and Mike Williams. 2022 · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems . Curran Associates Inc., Red Hook, NY, USA
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
Root Mean Square Layer Normalization . Curran Associates Inc., Red Hook, NY, USA
Biao Zhang and Rico Sennrich. 2019 · 2019
Cited alongside, same era.
Template-Based Question Generation from Retrieved Sentences for Improved Unsupervised Question Answering
Alexander Fabbri, Patrick Ng, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang. 2020 · 2020
Cited alongside, same era.
Prosodic Bootstrapping
Judit Gervain, Anne Christophe, and Reiko Mazuka. 2020 · 2020
Cited alongside, same era.
Multi-agent communication meets natural language: Synergies between functional and structural language learning
Angeliki Lazaridou, Anna Potapenko, and Olivier Tieleman. 2020 · 2020
Cited alongside, same era.
How Can We Accelerate Progress Towards Human-like Linguistic Generalization?
Tal Linzen. 2020 · 2020
Cited alongside, same era.
Internal and external pressures on language emergence: least effort, object constancy and frequency
Diana Rodríguez Luna, Edoardo Maria Ponti, Dieuwke Hupkes, and Elia Bruni. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2022 · 2022
Later among the works it cites.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind. 2022 · 2022
Later among the works it cites.
Linking emergent and natural languages via corpus transfer
Shunyu Yao, Mo Yu, Yang Zhang, Karthik R Narasimhan, Joshua B Tenenbaum, and Chuang Gan. 2022 · 2022
Later among the works it cites.
Putting Natural in Natural Language Processing
Grzegorz Chrupała. 2023 · 2023
Closest in time.
Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2023 · 2023
Closest in time.
The minipile challenge for data-efficient language models
Jean Kaddour. 2023 · 2023
Closest in time.
Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators
Andreas Liesenfeld, Alianda Lopez, and Mark Dingemanse. 2023 · 2023
Closest in time.
Omnigrok: Grokking beyond algorithmic data
Ziming Liu, Eric J Michaud, and Max Tegmark. 2023 · 2023
Closest in time.
Prosodic cues enhance infants’ sensitivity to nonadjacent regularities
Anna Martinez-Alvarez, Judit Gervain, Elena Koulaguina, Ferran Pons, and Ruth de Diego-Balaguer. 2023 · 2023
Closest in time.
Modeling rapid language learning by distilling Bayesian priors into artificial neural networks
R. Thomas McCoy and Thomas L. Griffiths. 2023 · 2023
Closest in time.
Inverse scaling: When bigger isn’t better
Ian R. McKenzie, Alexander Lyzhov, Michael Pieler, Alicia Parrish, Aaron Mueller, Ameya Prabhu, Euan McLean, Aaron Kirtland, Alexis Ross, Alisa Liu, Andrew Gritsevskiy, Daniel Wurgaft, Derik Kauffman, Gabriel Recchia, Jiacheng Liu, Joe Cavanagh, Max Weiss, Sicong Huang, The Floating Droid, Tom Tseng, Tomasz Korbak, Xudong Shen, Yuhui Zhang, Zhengping Zhou, Najoung Kim, Samuel R. Bowman, and Ethan Perez. 2023 · 2023
Closest in time.
Grokking of hierarchical structure in vanilla transformers
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher Manning. 2023 · 2023
Closest in time.
Pretrain on just structure: Understanding linguistic inductive biases using transfer learning
Isabel Papadimitriou and Dan Jurafsky. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Closest in time.
Quantifying the perceptual value of lexical and non-lexical channels in speech
Sarenne Wallbridge, Peter Bell, and Catherine Lai. 2023 · 2023
Closest in time.
Findings of the BabyLM Challenge: Sample-efficient pretraining on developmentally plausible corpora
Alex Warstadt, Leshem Choshen, Ryan Cotterell, Tal Linzen, Aaron Mueller, Ethan Wilcox, Williams Adina, and Chengxu Zhuang. 2023 · 2023
Closest in time.