COMET-22: Unbabel-IST 2022 submission for the metrics shared task
Rei, Ricardo, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. 2022 · 2022
Later among the works it cites.
Finding skill neurons in pre-trained transformer-based language models
Wang, Xiaozhi, Kaiyue Wen, Zhengyan Zhang, Lei Hou, Zhiyuan Liu, and Juanzi Li. 2022 · 2022
Later among the works it cites.
What artificial neural networks can tell us about human language acquisition
Original
Warstadt, Alex and Samuel R. Bowman. 2022 · 2022
Later among the works it cites.
The belebele benchmark: a parallel reading comprehension dataset in 122 language variants
Original
Bandarkar, Lucas, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2023 · 2023
Later among the works it cites.
On the proper role of linguistically oriented deep net analysis in linguistic theorising
Baroni, Marco. 2023 · 2023
Later among the works it cites.
Worldsense: A synthetic benchmark for grounded reasoning in large language models
Original
Benchekroun, Youssef, Megi Dervishi, Mark Ibrahim, Jean-Baptiste Gaya, Xavier Martinet, Grégoire Mialon, Thomas Scialom, Emmanuel Dupoux, Dieuwke Hupkes, and Pascal Vincent. 2023 · 2023
Later among the works it cites.
The reversal curse: LLMs trained on ”A is B” fail to learn ”B is A”
Original
Berglund, Lukas, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. 2023 · 2023
Later among the works it cites.
Chatgpt broke the turing test — the race is on for new ways to assess ai
Biever, Celeste. 2023 · 2023
Later among the works it cites.
Zero-shot approach to overcome perturbation sensitivity of prompts
Chakraborty, Mohna, Adithya Kulkarni, and Qi Li. 2023 · 2023
Later among the works it cites.
Language model behavior: A comprehensive survey
Original
Chang, Tyler A. and Benjamin K. Bergen. 2023 · 2023
Later among the works it cites.
Large language models demonstrate the potential of statistical learning in language
Contreras Kallens, Pablo, Ross Deans Kristensen-McLachlan, and Morten H Christiansen. 2023 · 2023
Later among the works it cites.
Shortcut learning of large language models in natural language understanding
Original
Du, Mengnan, Fengxiang He, Na Zou, Dacheng Tao, and Xia Hu. 2023 · 2023
Later among the works it cites.
The effect of scaling, retrieval augmentation and form on the factual consistency of language models
Hagström, Lovisa, Denitsa Saynova, Tobias Norlund, Moa Johansson, and Richard Johansson. 2023 · 2023
Later among the works it cites.
Rethinking reasoning evaluation with theories of intelligence
Heineman, David. 2023 · 2023
Later among the works it cites.
Surprisal does not explain syntactic disambiguation difficulty: evidence from a large-scale benchmark
Huang, Kuan-Jung, Suhas Arehalli, Mari Kugemoto, Christian Muxica, Grusha Prasad, Brian Dillon, and Tal Linzen. 2023 · 2023
Later among the works it cites.
State-of-the-art generalisation research in NLP: A taxonomy and review
Original
Hupkes, Dieuwke, Mario Giulianelli, Verna Dankers, Mikel Artetxe, Yanai Elazar, Tiago Pimentel, Christos Christodoulopoulos, Karim Lasri, Naomi Saphra, Arabella Sinclair, Dennis Ulmer, Florian Schottmann, Khuyagbaatar Batsuren, Kaiser Sun, Koustuv Sinha, Leila Khalatbari, Maria Ryskina, Rita Frieske, Ryan Cotterell, and Zhijing Jin. 2023 · 2023
Later among the works it cites.
Atlas: Few-shot learning with retrieval augmented language models
Izacard, Gautier, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023 · 2023
Later among the works it cites.
Consistency analysis of ChatGPT
Original
Jang, Myeongjun and Thomas Lukasiewicz. 2023 · 2023
Later among the works it cites.
What should replace the Turing Test?
Johnson-Laird, Philip N. and Marco Ragni. 2023 · 2023
Later among the works it cites.
The defeat of the Winograd Schema Challenge
Original
Kocijan, Vid, Ernest Davis, Thomas Lukasiewicz, Gary Marcus, and Leora Morgenstern. 2023 · 2023
Later among the works it cites.
Holistic evaluation of language models
Liang, Percy, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Alexander Cosgrove, Christopher D Manning, Christopher Re, Diana Acosta-Navas, Drew Arad Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue WANG, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Andrew Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2023 · 2023
Later among the works it cites.
Dissociating language and thought in large language models
Original
Mahowald, Kyle, Anna A. Ivanova, Idan A. Blank, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. 2023 · 2023
Later among the works it cites.
Do language models refer?
Original
Mandelkern, Matthew and Tal Linzen. 2023 · 2023
Later among the works it cites.
Embers of autoregression: Understanding large language models through the problem they are trained to solve
Original
McCoy, R. Thomas, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L. Griffiths. 2023 · 2023
Later among the works it cites.
Sources of hallucination by large language models on inference tasks
Original
McKenna, Nick, Tianyi Li, Liang Cheng, Mohammad Javad Hosseini, Mark Johnson, and Mark Steedman. 2023 · 2023
Later among the works it cites.
The debate over understanding in AI’s large language models
Mitchell, Melanie and David C. Krakauer. 2023 · 2023
Later among the works it cites.
State of what art? a call for multi-prompt llm evaluation
Original
Mizrahi, Moran, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. 2023 · 2023
Later among the works it cites.
The vector grounding problem
Original
Mollo, Dimitri Coelho and Raphaël Millière. 2023 · 2023
Later among the works it cites.
Separating form and meaning: Using self-consistency to quantify task understanding across multiple senses
Ohmer, Xenia, Elia Bruni, and Dieuwke Hupkes. 2023 · 2023
Later among the works it cites.
Symbols and grounding in large language models
Pavlick, Ellie. 2023 · 2023
Later among the works it cites.
Modern language models refute Chomsky’s approach to language
Piantadosi, Steven. 2023 · 2023
Later among the works it cites.
Cross-lingual consistency of factual knowledge in multilingual language models
Original
Qi, Jirui, Raquel Fernández, and Arianna Bisazza. 2023 · 2023
Later among the works it cites.
PECO: Examining single sentence label leakage in natural language inference datasets through progressive evaluation of cluster outliers
Saxon, Michael, Xinyi Wang, Wenda Xu, and William Yang Wang. 2023 · 2023
Later among the works it cites.
The validity of evaluation results: Assessing concurrence across compositionality benchmarks
Sun, Kaiser, Adina Williams, and Dieuwke Hupkes. 2023 · 2023
Later among the works it cites.
A language model with limited memory capacity captures interference in human sentence processing
Original
Timkey, William and Tal Linzen. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Original
Touvron, Hugo, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Later among the works it cites.
Are large language models really robust to word-level perturbations?
Original
Wang, Haoyu, Guozheng Ma, Cong Yu, Ning Gui, Linrui Zhang, Zhiqi Huang, Suwei Ma, Yongzhe Chang, Sen Zhang, Li Shen, Xueqian Wang, Peilin Zhao, and Dacheng Tao. 2023 · 2023
Later among the works it cites.
Call for papers – The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
Original
Warstadt, Alex, Leshem Choshen, Aaron Mueller, Adina Williams, Ethan Wilcox, and Chengxu Zhuang. 2023 · 2023
Later among the works it cites.
Mind the instructions: a holistic evaluation of consistency and interactions in prompt-based learning
Weber, Lucas, Elia Bruni, and Dieuwke Hupkes. 2023 · 2023
Later among the works it cites.