Fetching the paper…
Reading the bibliography…
Large Language Model (LLM) evaluation is currently one of the most important areas of research, with existing benchmarks proving to be insufficient and not completely representative of LLMs' various capabilities.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryściński, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 1910
Earlier work this paper cites.
A richly annotated corpus for different tasks in automated fact-checking
Andreas Hanselowski, Christian Stab, Claudia Schulz, Zile Li, and Iryna Gurevych. 2019 · 1911
Earlier work this paper cites.
Belief in conspiracy theories
Ted Goertzel. 1994 · 1994
Earlier work this paper cites.
Conspiracy theories
Cass R Sunstein and Adrian Vermeule. 2008 · 2008
Earlier work this paper cites.
Unanswered questions: A preliminary investigation of personality and individual difference predictors of 9/11 conspiracist beliefs
Viren Swami, Tomas Chamorro-Premuzic, and Adrian Furnham. 2010 · 2010
Earlier work this paper cites.
Measuring belief in conspiracy theories: The generic conspiracist beliefs scale
Robert Brotherton, Christopher C French, and Alan D Pickering. 2013 · 2013
Earlier work this paper cites.
Commercial conspiracy theories: A pilot study
Adrian Furnham. 2013 · 2013
Earlier work this paper cites.
The measurement and prediction of conspiracy beliefs
Chelsea Rose. 2017 · 2017
Earlier work this paper cites.
" liar, liar pants on fire": A new benchmark dataset for fake news detection
William Yang Wang. 2017 · 2017
Earlier work this paper cites.
Where is your evidence: Improving fact-checking by justification modeling
Tariq Alhindi, Savvas Petridis, and Smaranda Muresan. 2018 · 2018
Earlier work this paper cites.
Universal sentence encoder for English
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil. 2018 · 2018
Cited alongside, same era.
Negligent falsehood, white ignorance, and false news
Jennifer Saul, E Michaelson, and A Stokke. 2018 · 2018
Cited alongside, same era.
Belief in conspiracy theories: Basic principles of an emerging research domain
Jan-Willem van Prooijen and Karen M Douglas. 2018 · 2018
Cited alongside, same era.
Connecting the dots: Illusory pattern perception predicts belief in conspiracies and the supernatural
Jan-Willem Van Prooijen, Karen M Douglas, and Clara De Inocencio. 2018 · 2018
Cited alongside, same era.
Increased conspiracy beliefs among ethnic and muslim minorities
Jan-Willem van Prooijen, Jaap Staman, and André PM Krouwel. 2018 · 2018
Cited alongside, same era.
Assessing the factual accuracy of generated text
Finding someone to blame: The link between covid-19 conspiracy beliefs, prejudice, support for violence, and other negative social outcomes
Jakub Šrol, Vladimíra Čavojová, and Eva Ballová Mikušková. 2022 · 2022
Later among the works it cites.
Evaluating the factual consistency of large language models through summarization
Derek Tam, Anisha Mascarenhas, Shiyue Zhang, Sarah Kwan, Mohit Bansal, and Colin Raffel. 2022 · 2022
Later among the works it cites.
Emergent abilities of large language models
Barret Zoph, Colin Raffel, Dale Schuurmans, Dani Yogatama, Denny Zhou, Don Metzler, Ed H. Chi, Jason Wei, Jeff Dean, Liam B. Fedus, Maarten Paul Bosma, Oriol Vinyals, Percy Liang, Sebastian Borgeaud, Tatsunori B. Hashimoto, and Yi Tay. 2022 · 2022
Later among the works it cites.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Baber Abbasi, et al. 2023 · 2023
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ben Goodrich, Vinay Rao, Peter J Liu, and Mohammad Saleh. 2019 · 2019
Cited alongside, same era.
Checkthat! at clef 2020: Enabling the automatic identification and verification of claims in social media
Alberto Barrón-Cedeno, Tamer Elsayed, Preslav Nakov, Giovanni Da San Martino, Maram Hasanain, Reem Suwaileh, and Fatima Haouari. 2020 · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2021 · 2021
Cited alongside, same era.
Entity-level factual consistency of abstractive text summarization
Feng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, and Bing Xiang. 2021a
Cited in the paper.
Improving factual consistency of abstractive summarization via question answering
Feng Nan, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O Arnold, and Bing Xiang. 2021b
Cited in the paper.
Later among the works it cites.
Reliability check: An analysis of GPT-3’s response to sensitive topics and prompt wording
Aisha Khatun and Daniel Brown. 2023 · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, and et al. 2023 · 2023
Later among the works it cites.
50 fox news ’lies’ in 6 seconds, from ’the daily show’
Lauren Carroll and Aaron Sharockman. 2015 · 2024
Closest in time.
Rag is just fancier prompt engineering
Mohit Pandey. 2023 · 2024
Closest in time.