Fetching the paper…
Reading the bibliography…
NLP benchmarks rely on standardized datasets for training and evaluating models and are crucial for advancing the field.
Self: Learning to filter noisy labels with self-ensembling
Duc Tam Nguyen, Chaithanya Kumar Mummadi, Thi-Phuong-Nhung Ngo, Thi Hoai Phuong Nguyen, Laura Beggel, and Thomas Brox. 2019 · 1910
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan T. McDonald. 2020 · 1919
Earlier work this paper cites.
The use of confidence or fiducial limits illustrated in the case of the binomial
C. J. Clopper and E. S. Pearson. 1934 · 1934
Earlier work this paper cites.
Estimating the reliability, systematic error, and random error of interval data
Klaus Krippendorff. 1970 · 1970
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
Joseph L. Fleiss. 1971 · 1971
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire. 1997 · 1997
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Identifying mislabeled data using the area under the margin ranking
Geoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, and Kilian Q. Weinberger. 2020 · 2001
Earlier work this paper cites.
Ensemble methods in machine learning
Thomas G. Dietterich. 2007 · 2007
Earlier work this paper cites.
Cheap and fast – but is it good? evaluating non-expert annotations for natural language tasks
Rion Snow, Brendan T. O’Connor, Dan Jurafsky, and A. Ng. 2008 · 2008
Earlier work this paper cites.
An extended model of natural logic
Bill MacCartney and Christopher D. Manning. 2009 · 2009
Earlier work this paper cites.
Learning to combine discriminative classifiers: confidence based
Chi-Hoon Lee. 2010 · 2010
Earlier work this paper cites.
Translational biomarker discovery in clinical metabolomics: an introductory tutorial
Jianguo Xia, David I. Broadhurst, Michael Wilson, and David Scott Wishart. 2012 · 2012
Earlier work this paper cites.
Quality control in crowdsourcing systems: Issues and directions
Mohammad Allahbakhsh, Boualem Benatallah, Aleksandar Ignjatovic, Hamid Reza Motahari-Nezhad, Elisa Bertino, and Schahram Dustdar. 2013 · 2013
Earlier work this paper cites.
An analysis of human factors and label accuracy in crowdsourcing relevance judgments
Gabriella Kazai, Jaap Kamps, and Natasa Milic-Frayling. 2013 · 2013
Earlier work this paper cites.
Investigating the disagreement between clinicians’ ratings of patients in icus
Simon Rogers, Derek H. Sleeman, and John Kinsella. 2013 · 2013
Earlier work this paper cites.
Classification in the presence of label noise: A survey
Benoît Frénay and Michel Verleysen. 2014 · 2014
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2016 · 2016
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cissé, Yann N. Dauphin, and David Lopez-Paz. 2018 · 2018
Earlier work this paper cites.
An MTurk crisis? shifts in data quality and the impact on study results
Michael Chmielewski and Sarah C. Kucker. 2019 · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019 · 2019
Earlier work this paper cites.
Confident learning: Estimating uncertainty in dataset labels
Curtis G. Northcutt, Lu Jiang, and Isaac L. Chuang. 2019 · 2019
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
PAWS: Paraphrase adversaries from word scrambling
Yuan Zhang, Jason Baldridge, and Luheng He. 2019 · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
Understanding the tradeoff between cost and quality of expert annotations for keyphrase extraction
Hung Chau, Saeid Balaneshin, Kai Liu, and Ondrej Linda. 2020 · 2020
Cited alongside, same era.
Inaccurate labels in weakly-supervised deep learning: Automatic identification and correction and their impact on classification performance
Degan Hao, Lei Zhang, Jules H. Sumkin, Aly A. Mohamed, and Shandong Wu. 2020 · 2020
Cited alongside, same era.
The shape of and solutions to the MTurk quality crisis
Ryan Kennedy, Scott Clifford, Tyler Burleigh, Philip D Waggoner, Ryan Jewell, and Nicholas JG Winter. 2020 · 2020
Cited alongside, same era.
Research on data quality control of crowdsourcing annotation: A survey
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Predictive performance of multi-model ensemble forecasts of covid-19 across european nations
Katharine Sherratt, Hugo Gruson, Rok Grah, Helen Johnson, Rene Niehus, Bastian Prasse, and et al. 2023 · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali Safaya, Ali Tazarv, Alice Xiang, Alicia Parrish, Allen Nie, Aman Hussain, Amanda Askell, Amanda Dsouza, Ambrose Slone, Ameet Rahane, Anantharaman S. Iyer, Anders Andreassen, and et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jian Lu, Wei Li, Qingren Wang, and Yiwen Zhang. 2020 · 2020
Cited alongside, same era.
Identifying incorrect labels in the conll-2003 corpus
Frederick Reiss, Hong Xu, Bryan Cutler, Karthik Muthuraman, and Zachary Eichenberger. 2020 · 2020
Cited alongside, same era.
Summeval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir R. Radev. 2021 · 2021
Cited alongside, same era.
Evaluating CloudResearch’s approved group as a solution for problematic data quality on MTurk
David N. Hauser, Aaron J. Moss, Cheskie Rosenzweig, Shalom N. Jaffe, Jonathan Robinson, and Leib Litman. 2021 · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
Pervasive label errors in test sets destabilize machine learning benchmarks
Curtis G. Northcutt, Anish Athalye, and Jonas W. Mueller. 2021 · 2021
Cited alongside, same era.
Get your vitamin C! robust fact verification with contrastive evidence
Tal Schuster, Adam Fisch, and Regina Barzilay. 2021 · 2021
Cited alongside, same era.
Learning from disagreement: A survey
Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021 · 2021
Cited alongside, same era.
With a little push, NLI models can robustly and efficiently predict faithfulness
Julius Steen, Juri Opitz, Anette Frank, and Katja Markert. 2023 · 2023
Later among the works it cites.
The impact of inconsistent human annotations on AI driven clinical decision making
Aneeta Sylolypavan, Derek H. Sleeman, Honghan Wu, and Malcolm Sim. 2023 · 2023
Later among the works it cites.
Evaluating the factual consistency of large language models through news summarization
Derek Tam, Anisha Mascarenhas, Shiyue Zhang, Sarah Kwan, Mohit Bansal, and Colin Raffel. 2023 · 2023
Later among the works it cites.
Petter Törnberg. 2023 · 2023
Later among the works it cites.
Navigating cultural chasms: Exploring and unlocking the cultural POV of text-to-image models
Mor Ventura, Eyal Ben-David, Anna Korhonen, and Roi Reichart. 2023 · 2023
Later among the works it cites.
ActiveAED: A human in the loop improves annotation error detection
Leon Weber and Barbara Plank. 2023 · 2023
Later among the works it cites.
Improving opinion-based question answering systems through label error detection and overwrite
Xiao Yang, Ahmed K. Mohamed, Shashank Jain, Stanislav Peshterliev, Debojeet Chatterjee, Hanwen Zha, Nikita Bhalla, Gagan Aneja, and Pranab Mohanty. 2023 · 2023
Later among the works it cites.
Alignscore: Evaluating factual consistency with a unified alignment function
Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu. 2023 · 2023
Later among the works it cites.
LLMaAA: Making large language models as active annotators
Ruoyu Zhang, Yanzeng Li, Yongliang Ma, Ming Zhou, and Lei Zou. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Later among the works it cites.
Measuring the robustness of nlp models to domain shifts
Nitay Calderon, Naveh Porat, Eyal Ben-David, Alexander Chapanin, Zorik Gekhman, Nadav Oved, Vitaly Shalumov, and Roi Reichart. 2024 · 2024
Closest in time.
On behalf of the stakeholders: Trends in NLP model interpretability in the era of llms
Nitay Calderon and Roi Reichart. 2024 · 2024
Closest in time.
Is a large language model a good annotator for event extraction?
Ruirui Chen, Chengwei Qin, Weifeng Jiang, and Dongkyu Choi. 2024 · 2024
Closest in time.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, and et al. 2024 · 2024
Closest in time.
Gpt is not an annotator: The necessity of human annotation in fairness benchmark construction
Virginia K. Felkner, Jennifer A. Thompson, and Jonathan May. 2024 · 2024
Closest in time.
Faithful explanations of black-box NLP models using LLM-generated counterfactuals
Yair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin, Amit Sharma, and Roi Reichart. 2024 · 2024
Closest in time.
Nataliia Kholodna, Sahib Julka, Mohammad Khodadadi, Muhammed Nurullah Gumus, and Michael Granitzer. 2024 · 2024
Closest in time.
Meganno+: A human-llm collaborative annotation system
Han Jun Kim, Kushan Mitra, Rafael Li Chen, Sajjadur Rahman, and Dan Zhang. 2024 · 2024
Closest in time.
Beyond sole strength: Customized ensembles for generalized vision-language models
Zhihe Lu, Jiawang Bai, Xin Li, Zeyu Xiao, and Xinchao Wang. 2024 · 2024
Closest in time.
Less is more for improving automatic evaluation of factual consistency
Tong Wang, Ninad Kulkarni, and Yanjun Qi. 2024 · 2024
Closest in time.
VariErr NLI: Separating annotation error from human label variation
Leon Weber-Genzel, Siyao Peng, Marie-Catherine De Marneffe, and Barbara Plank. 2024 · 2024
Closest in time.
LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, Andre Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K Surikuchi, Ece Takmaz, and Alberto Testoni. 2025 · 2025
Closest in time.
TrueTeacher: Learning factual consistency evaluation with large language models
Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. 2023 · 2070
Closest in time.
The colorful future of llms: Evaluating and improving llms as emotional supporters for queer youth
Shir Lissak, Nitay Calderon, Geva Shenkman, Yaakov Ophir, Eyal Fruchter, Anat Brunstein Klomek, and Roi Reichart. 2024 · 2079
Closest in time.