Fetching the paper…
Reading the bibliography…
Today's best language models still struggle with hallucinations: factually incorrect generations, which impede their ability to reliably retrieve information seen during training.
Clever hans : the horse of mr. von osten, 1911
Oskar Pfungst and Robert Rosenthal · 1911
Earlier work this paper cites.
Web based probabilistic textual entailment
Oren Glickman, Ido Dagan, and Moshe Koppel · 2005
Earlier work this paper cites.
Textual entailment through extended lexical overlap and lexico-semantic matching
Rod Adams, Gabriel Nicolae, Cristina Nicolae, and Sanda Harabagiu · 2007
Earlier work this paper cites.
Pre-training a language model without human language
David Cheng-Han Chiang and Hung-yi Lee · 2012
Earlier work this paper cites.
Stress test evaluation for natural language inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig · 2018
Earlier work this paper cites.
Behavior analysis of NLI models: Uncovering the influence of three factors on robustness
Ivan Sanchez, Jeff Mitchell, and Sebastian Riedel · 2018
Earlier work this paper cites.
Evaluating compositionality in sentence embeddings
Ishita Dasgupta, Demi Guo, Andreas Stuhlmüller, Samuel J Gershman, and Noah D Goodman · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Linking artificial and human neural representations of language
Jon Gauthier and Roger Levy · 2019
Earlier work this paper cites.
RoBERTa: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
CLUTRR: A diagnostic benchmark for inductive reasoning from text
Koustuv Sinha, Shagun Sodhani, Jin Dong, Joelle Pineau, and William L. Hamilton · 2019
Earlier work this paper cites.
GenWiki: A dataset of 1.3 million content-sharing text and graphs for unsupervised graph-to-text generation
Zhijing Jin, Qipeng Guo, Xipeng Qiu, and Zheng Zhang · 2020
Cited alongside, same era.
SpanBERT: Improving Pre-training by Representing and Predicting Spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy · 2020
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding, 2020
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 2020
Cited alongside, same era.
Masked language modeling and the distributional hypothesis: Order word matters pre-training for little
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways, 2022
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel · 2022
Mistral 7b, 2023
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
Do large language models need sensory ground- ing for meaning and understanding?, 2023
Yann LeCun · 2023
Later among the works it cites.
Are we falling in a middle-intelligence trap? an analysis and mitigation of the reversal curse, 2023
Ang Lv, Kaiyi Zhang, Shufang Xie, Quan Tu, Yuhan Chen, Ji-Rong Wen, and Rui Yan · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI · 2023
Later among the works it cites.
Should you mask 15% in masked language modeling?, 2023
Alexander Wettig, Tianyu Gao, Zexuan Zhong, and Danqi Chen · 2023
Later among the works it cites.
Physics of language models: Part 3.1, knowledge storage and extraction, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Prioritized training on points that are learnable, worth learning, and not yet learnt, 2022
Sören Mindermann, Jan Brauner, Muhammed Razzak, Mrinank Sharma, Andreas Kirsch, Winnie Xu, Benedikt Höltgen, Aidan N. Gomez, Adrien Morisot, Sebastian Farquhar, and Yarin Gal · 2022
Cited alongside, same era.
Looking at the overlooked: An analysis on the word-overlap bias in natural language inference
Sara Rajaee, Yadollah Yaghoobzadeh, and Mohammad Taher Pilehvar · 2022
Cited alongside, same era.
Ul2: Unifying language learning paradigms
Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier García, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler · 2022
Cited alongside, same era.
ANLIzing the adversarial natural language inference dataset
Adina Williams, Tristan Thrush, and Douwe Kiela · 2022
Cited alongside, same era.
Physics of language models: Part 3.2, knowledge manipulation, 2023
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Cited alongside, same era.
Structured denoising diffusion models in discrete state-spaces, 2023
Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg · 2023
Cited alongside, same era.
The reversal curse: Llms trained on ”a is b” fail to learn ”b is a”, 2023
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans · 2023
Cited alongside, same era.
Zeyuan Allen Zhu and Yuanzhi Li · 2023
Later among the works it cites.
The pitfalls of next-token prediction, 2024
Gregor Bachmann and Vaishnavh Nagarajan · 2024
Closest in time.
Hallucinating Law: Legal Mistakes with Large Language Models are Pervasive, 2024
Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E. Ho · 2024
Closest in time.
Better & faster large language models via multi-token prediction, 2024
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, and Gabriel Synnaeve · 2024
Closest in time.
Reverse training to nurse the reversal curse, 2024
Olga Golovneva, Zeyuan Allen-Zhu, Jason Weston, and Sainbayar Sukhbaatar · 2024
Closest in time.
Disk: A diffusion model for structured knowledge, 2024
Ouail Kitouni, Niklas Nolte, James Hensman, and Bhaskar Mitra · 2024
Closest in time.
Memory mosaics, 2024
Jianyu Zhang, Niklas Nolte, Ranajoy Sadhukhan, Beidi Chen, and Léon Bottou · 2024
Closest in time.