The dangers of underclaiming: Reasons for caution when reporting how NLP systems fail
Samuel Bowman. 2022 · 2022
Closest in time.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022 · 2022
Closest in time.
Language models show human-like content effects on reasoning
Original
Ishita Dasgupta, Andrew K Lampinen, Stephanie CY Chan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill. 2022 · 2022
Closest in time.
Experimentology: An open science approach to experimental psychology methods
Michael C Frank, Mika Braginsky, Julie Cachia, Nicholas Coles, Tom Hardwicke, Robert Hawkins, Maya Mathur, and Rondeline Williams. 2022 · 2022
Closest in time.
A resource-rational model of human processing of recursive linguistic structure
Michael Hahn, Richard Futrell, Roger Levy, and Edward Gibson. 2022 · 2022
Closest in time.
Training compute-optimal large language models
Original
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022 · 2022
Closest in time.
Artificial neural network language models align neurally and behaviorally with humans even after a developmentally realistic amount of training
Eghbal A Hosseini, Martin A Schrimpf, Yian Zhang, Samuel Bowman, Noga Zaslavsky, and Evelina Fedorenko. 2022 · 2022
Closest in time.
Large language models are zero-shot reasoners
Original
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Closest in time.
Can transformers process recursive nested constructions, like humans?
Yair Lakretz, Théo Desbordes, Dieuwke Hupkes, and Stanislas Dehaene. 2022 · 2022
Closest in time.
Can language models learn from explanations in context?
Original
Andrew K Lampinen, Ishita Dasgupta, Stephanie CY Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L McClelland, Jane X Wang, and Felix Hill. 2022 · 2022
Closest in time.
Comps: Conceptual minimal pair sentences for testing property knowledge and inheritance in pre-trained language models
Original
Kanishka Misra, Julia Taylor Rayz, and Allyson Ettinger. 2022 · 2022
Closest in time.
Training language models to follow instructions with human feedback
Original
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Closest in time.
When a sentence does not introduce a discourse entity, transformer-based models still sometimes refer to it
Original
Sebastian Schuster and Tal Linzen. 2022 · 2022
Closest in time.
Prompting gpt-3 to be reliable
Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, and Lijuan Wang. 2022 · 2022
Closest in time.
Structural persistence in language models: Priming as a window into abstract language representations
Arabella Sinclair, Jaap Jumelet, Willem Zuidema, and Raquel Fernández. 2022 · 2022
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Original
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022 · 2022
Closest in time.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. 2022 · 2022
Closest in time.
Transcending scaling laws with 0.1% extra compute
Original
Yi Tay, Jason Wei, Hyung Won Chung, Vinh Q Tran, David R So, Siamak Shakeri, Xavier Garcia, Huaixiu Steven Zheng, Jinfeng Rao, Aakanksha Chowdhery, et al. 2022 · 2022
Closest in time.
Large language models still can’t plan (a benchmark for llms on planning and reasoning about change)
Original
Karthik Valmeekam, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati. 2022 · 2022
Closest in time.
What language model architecture and pretraining objective work best for zero-shot generalization?
Original
Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel. 2022 · 2022
Closest in time.
The generalizability crisis
Tal Yarkoni. 2022 · 2022
Closest in time.
Least-to-most prompting enables complex reasoning in large language models
Original
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2022 · 2022
Closest in time.
Are language models worse than humans at following prompts? it’s complicated
Original
Albert Webson, Alyssa Marie Loo, Qinan Yu, and Ellie Pavlick. 2023 · 2023
Closest in time.