Fetching the paper…
Reading the bibliography…
This paper introduces a framework for the automated evaluation of natural language texts.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
CTRL: A conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav Varshney, Caiming Xiong, and Richard Socher. 2019 · 1909
Earlier work this paper cites.
Stop measuring calibration when humans disagree
Joris Baan, Wilker Aziz, Barbara Plank, and Raquel Fernandez. 2022 · 1915
Earlier work this paper cites.
Multicalibration: Calibration for the (computationally-identifiable) masses
Úrsula Hébert-Johnson, Michael Kim, Omer Reingold, and Guy Rothblum. 2018 · 1948
Earlier work this paper cites.
The use of the computer in analyzing student essays
E. B. Page. 1968 · 1968
Earlier work this paper cites.
Overview of the Fourth Text REtrieval Conference (TREC-4)
Donna K. Harman. 1996 · 1996
Earlier work this paper cites.
Nonlinear dimensionality reduction by locally linear embedding
S. T. Roweis and L. K. Saul. 2000 · 2000
Earlier work this paper cites.
A global geometric framework for nonlinear dimensionality reduction
J. B. Tenenbaum, V. D. Silva, and J. C. Langford. 2000 · 2000
Earlier work this paper cites.
The AMI meeting corpus: A pre-announcement
Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, Guillaume Lathoud, Mike Lincoln, Agnes Lisowska, Iain McCowan, Wilfried Post, Dennis Reidsma, and Pierre Wellner. 2005 · 2005
Earlier work this paper cites.
Predicting good probabilities with supervised learning
Alexandru Niculescu-Mizil and Rich Caruana. 2005 · 2005
Earlier work this paper cites.
Cost-sensitive dynamic feature selection
He He, Hal Daumé III, and Jason Eisner. 2012 · 2012
Earlier work this paper cites.
Active Learning
Burr Settles. 2012 · 2012
Earlier work this paper cites.
The Toronto Paper Matching System: An automated paper-reviewer assignment system
Laurent Charlin and Richard S. Zemel. 2013 · 2013
Earlier work this paper cites.
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty. 2015 · 2015
Earlier work this paper cites.
The Ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems
Ryan Lowe, Nissan Pow, Iulian Serban, and Joelle Pineau. 2015 · 2015
Earlier work this paper cites.
Reducing click and skip errors in search result ranking
Jiepu Jiang and James Allan. 2016 · 2016
Earlier work this paper cites.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Earlier work this paper cites.
Neural autoregressive distribution estimation
Benigno Uria, Marc-Alexandre Côté, Karol Gregor, Iain Murray, and Hugo Larochelle. 2016 · 2016
Earlier work this paper cites.
Do not use averages with Likert scale data
Dwight Barry. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Dynamic weights in multi-objective deep reinforcement learning
Axel Abels, Diederik Roijers, Tom Lenaerts, Ann Nowé, and Denis Steckelmacher. 2019 · 2019
Earlier work this paper cites.
Dynamic feature acquisition using denoising autoencoders
Mohammad Kachuee, Sajad Darabi, Babak Moatamed, and Majid Sarrafzadeh. 2019 · 2019
Earlier work this paper cites.
An introduction to variational autoencoders
Diederik P. Kingma and Max Welling. 2019 · 2019
Earlier work this paper cites.
Inherent disagreements in human textual inferences
Ellie Pavlick and Tom Kwiatkowski. 2019 · 2019
Earlier work this paper cites.
Technology Assisted Review (TAR) guidelines
Mike Quartararo, Matt Poplawski, Adam Strayer, et al. 2019 · 2019
Earlier work this paper cites.
Controllable neural story plot generation via reward shaping
Pradyumna Tambwekar, Murtaza Dhuliawala, Lara J. Martin, Animesh Mehta, Brent Harrison, and Mark O. Riedl. 2019 · 2019
Earlier work this paper cites.
A generalized algorithm for multi-objective reinforcement learning and policy adaptation
Runzhe Yang, Xingyuan Sun, and Karthik Narasimhan. 2019 · 2019
Earlier work this paper cites.
Methods in predictive techniques for mental health status on social media: A critical review
Stevie Chancellor and Munmun De Choudhury. 2020 · 2020
Cited alongside, same era.
Natural language inference with mixed effects
William Gantt, Benjamin Kane, and Aaron Steven White. 2020 · 2020
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020 · 2020
Cited alongside, same era.
We need to consider disagreement in evaluation
Valerio Basile, Michael Fell, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio, and Alexandra Uma. 2021 · 2021
Cited alongside, same era.
SemEval-2021 task 12: Learning with disagreements
Alexandra Uma, Tommaso Fornaciari, Anca Dumitrache, Tristan Miller, Jon Chamberlain, Barbara Plank, Edwin Simpson, and Massimo Poesio. 2021a · 2021
Cited alongside, same era.
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Later among the works it cites.
Rewarded soups: Towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Alexandre Ramé, Guillaume Couairon, Mustafa Shukor, Corentin Dancette, Jean-Baptiste Gaya, Laure Soulier, and Matthieu Cord. 2023 · 2023
Later among the works it cites.
ARES: An automated evaluation framework for retrieval-augmented generation systems
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2023 · 2023
Later among the works it cites.
Why don’t you do it right? analysing annotators’ disagreement in subjective tasks
Marta Sandri, Elisa Leonardelli, Sara Tonelli, and Elisabetta Jezek. 2023 · 2023
Later among the works it cites.
First tragedy, then parse: History repeats itself in the new era of large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human evaluation of automatically generated text: Current trends and best practice guidelines
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, and Emiel Krahmer. 2021 · 2021
Cited alongside, same era.
Tell me what you read: Automatic expertise-based annotator assignment for text annotation in expert domains
Hiyori Yoshikawa, Tomoya Iwakura, Kimi Kaneko, Hiroaki Yoshida, Yasutaka Kumano, Kazutaka Shimada, Rafal Rzepka, and Patrycja Swieczkowska. 2021 · 2021
Cited alongside, same era.
Natural language model for automatic identification of intimate partner violence reports from Twitter
Mohammed Ali Al-Garadi, Sangmi Kim, Yuting Guo, Elise Warren, Yuan-Chi Yang, Sahithi Lakamana, and Abeed Sarker. 2022 · 2022
Cited alongside, same era.
Constitutional AI: Harmlessness from AI feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022 · 2022
Cited alongside, same era.
Decomposing and recomposing event structure
William Gantt, Lelia Glass, and Aaron Steven White. 2022 · 2022
Cited alongside, same era.
On embeddings for numerical features in tabular deep learning
Yury Gorishniy, Ivan Rubachev, and Artem Babenko. 2022 · 2022
Cited alongside, same era.
ClueWeb22: 10 billion web documents with rich information
Arnold Overwijk, Chenyan Xiong, and Jamie Callan. 2022 · 2022
Cited alongside, same era.
Naomi Saphra, Eve Fleisig, Kyunghyun Cho, and Adam Lopez. 2023 · 2023
Later among the works it cites.
Auxiliary learning as an asymmetric bargaining game
Aviv Shamsian, Aviv Navon, Neta Glazer, Kenji Kawaguchi, Gal Chechik, and Ethan Fetaya. 2023 · 2023
Later among the works it cites.
Large language models enable few-shot clustering
Vijay Viswanathan, Kiril Gashteovski, Carolin Lawrence, Tongshuang Wu, and Graham Neubig. 2023 · 2023
Later among the works it cites.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2023 · 2023
Later among the works it cites.
Evaluating evaluation metrics: A framework for analyzing NLG evaluation metrics using measurement theory
Ziang Xiao, Susu Zhang, Vivian Lai, and Q. Vera Liao. 2023 · 2023
Later among the works it cites.
Automatic evaluation of attribution by large language models
Xiang Yue, Boshi Wang, Ziru Chen, Kai Zhang, Yu Su, and Huan Sun. 2023 · 2023
Later among the works it cites.
Conversational information seeking
Hamed Zamani, Johanne R. Trippas, Jeff Dalton, and Filip Radlinski. 2023 · 2023
Later among the works it cites.
ClusterLLM: Large language models as a guide for text clustering
Yuwei Zhang, Zihan Wang, and Jingbo Shang. 2023 · 2023
Later among the works it cites.
Theodore Zhao, Mu Wei, J. Samuel Preston, and Hoifung Poon. 2023 · 2023
Later among the works it cites.
Non-programmers can label programs indirectly via active examples: A case study with text-to-SQL
Ruiqi Zhong, Charlie Snell, Dan Klein, and Jason Eisner. 2023 · 2023
Later among the works it cites.
MEGAVERSE: Benchmarking large language models across languages, modalities, models and tasks
Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Prachi Jain, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2024 · 2024
Closest in time.
FACT-GPT: Fact-checking augmentation via claim matching with LLMs
Eun Cheol Choi and Emilio Ferrara. 2024 · 2024
Closest in time.
Multicalibration for confidence scoring in LLMs
Gianluca Detommaso, Martin Bertran, Riccardo Fogliato, and Aaron Roth. 2024 · 2024
Closest in time.
Assessment of a Large Language Model’s Responses to Questions and Cases About Glaucoma and Retina Management
Andy S. Huang, Kyle Hirabayashi, Laura Barna, Deep Parikh, and Louis R. Pasquale. 2024 · 2024
Closest in time.
Interpretable user satisfaction estimation for conversational systems with large language models
Ying-Chun Lin, Jennifer Neville, Jack W. Stokes, Longqi Yang, Tara Safavi, Mengting Wan, Scott Counts, Siddharth Suri, Reid Andersen, Xiaofeng Xu, Deepak Gupta, Sujay Kumar Jauhar, Xia Song, Georg Buscher, Saurabh Tiwary, Brent Hecht, and Jaime Teevan. 2024 · 2024
Closest in time.
Do AIs know what the most important issue is? using language models to code open-text social survey responses at scale
Jonathan Mellon, Jack Bailey, Ralph Scott, James Breckwoldt, Marta Miori, and Phillip Schmedeman. 2024 · 2024
Closest in time.
Using LLMs to bring evidence-based feedback into the classroom: AI-generated feedback increases secondary students’ text revision, motivation, and positive emotions
Jennifer Meyer, Thorben Jansen, Ronja Schiller, Lucas W. Liebenow, Marlene Steinbach, Andrea Horbach, and Johanna Fleckenstein. 2024 · 2024
Closest in time.
OpenAI GPT-3.5 Turbo 16K [ gpt-3.5-turbo-16k-0613
OpenAI. 2024 · 2024
Closest in time.
Large language models can accurately predict searcher preferences
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024 · 2024
Closest in time.
Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge, Furu Wei, and Heng Ji. 2024 · 2024
Closest in time.
Enhancing systematic decompositional natural language inference using informal logic
Nathaniel Weir, Kate Sanders, Orion Weller, Shreya Sharma, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Jansen, Peter Clark, and Benjamin Van Durme. 2024 · 2024
Closest in time.
Mental-LLM: Leveraging large language models for mental health prediction via online text data
Xuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel, Hong Yu, James Hendler, Marzyeh Ghassemi, Anind K. Dey, and Dakuo Wang. 2024 · 2024
Closest in time.
TrueTeacher: Learning factual consistency evaluation with large language models
Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. 2023 · 2070
Closest in time.