Quantifying appropriateness of summarization data for curriculum learning
Ryuji Kano, Takumi Takahashi, Toru Nishino, Motoki Taniguchi, Tomoki Taniguchi, and Tomoko Ohkuma. 2021 · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
Original
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario vSavsko, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clement Delangue, Th’eo Matussiere, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, Franccois Lagunas, Alexander M. Rush, and Thomas Wolf. 2021 · 2021
Later among the works it cites.
Subjective bias in abstractive summarization
Original
Lei Li, Wei Liu, Marina Litvak, Natalia Vanetik, Jiacheng Pei, Yinan Liu, and Siya Qi. 2021 · 2021
Later among the works it cites.
Are we learning yet? a meta review of evaluation failures across machine learning
Thomas Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt. 2021 · 2021
Later among the works it cites.
Scientific credibility of machine translation research: A meta-evaluation of 769 papers
Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021 · 2021
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
Original
Curtis G. Northcutt, Anish Athalye, and Jonas Mueller. 2021 · 2021
Later among the works it cites.
Data and its (dis)contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily L. Denton, and A. Hanna. 2021 · 2021
Later among the works it cites.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021 · 2021
Later among the works it cites.
Ai and the everything in the whole wide world benchmark
Original
Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada, Emily L. Denton, and A. Hanna. 2021 · 2021
Later among the works it cites.
Evaluation examples are not equally informative: How should that change NLP leaderboards?
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, and Jordan Boyd-Graber. 2021 · 2021
Later among the works it cites.
Challenges and opportunities in nlp benchmarking
Sebastian Ruder. 2021 · 2021
Later among the works it cites.
Targeting the benchmark: On methodology in current natural language processing research
David Schlangen. 2021 · 2021
Later among the works it cites.
Including signed languages in natural language processing
Kayo Yin, Amit Moryossef, Julie Hochgesang, Yoav Goldberg, and Malihe Alikhani. 2021 · 2021
Later among the works it cites.
Are larger pretrained language models uniformly better? comparing performance at the instance level
Ruiqi Zhong, Dhruba Ghosh, Dan Klein, and Jacob Steinhardt. 2021 · 2021
Later among the works it cites.
Aces: Translation accuracy challenge sets for evaluating machine translation metrics
Original
Chantal Amrhein, Nikita Moghe, and Liane Guillou. 2022 · 2022
Later among the works it cites.
Stop measuring calibration when humans disagree
Original
Joris Baan, Wilker Aziz, Barbara Plank, and Raquel Fernandez. 2022 · 2022
Later among the works it cites.
Making intelligence: Ethics, iq, and ml benchmarks
Original
Borhane Blili-Hamelin and Leif Hancox-Li. 2022 · 2022
Later among the works it cites.
Evaluation for change
Original
Rishi Bommasani. 2022 · 2022
Later among the works it cites.
What’s different between visual question answering for machine “understanding” versus for accessibility?
Yang Trista Cao, Kyle Seelman, Kyungjun Lee, and Hal Daum’e. 2022 · 2022
Later among the works it cites.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Aida Mostafazadeh Davani, Mark D’iaz, and Vinodkumar Prabhakaran. 2022 · 2022
Later among the works it cites.
Re-examining system-level correlations of automatic summarization evaluation metrics
Original
Daniel Deutsch, Rotem Dror, and Dan Roth. 2022 · 2022
Later among the works it cites.
On the origin of hallucinations in conversational models: Is it the datasets or the models?
Original
Nouha Dziri, Sivan Milton, Mo Yu, Osmar R Zaiane, and Siva Reddy. 2022 · 2022
Later among the works it cites.
Finding dataset shortcuts with grammar induction
Original
Dan Friedman, Alexander Wettig, and Danqi Chen. 2022 · 2022
Later among the works it cites.
An analysis of negation in natural language understanding corpora
Md Mosharaf Hossain, Dhivya Chinnappa, and Eduardo Blanco. 2022 · 2022
Later among the works it cites.
State-of-the-art generalisation research in nlp: a taxonomy and review
Original
Dieuwke Hupkes, Mario Giulianelli, Verna Dankers, Mikel Artetxe, Yanai Elazar, Tiago Pimentel, Christos Christodoulopoulos, Karim Lasri, Naomi Saphra, Arabella J. Sinclair, Dennis Ulmer, Florian Schottmann, Khuyagbaatar Batsuren, Kaiser Sun, Koustuv Sinha, Leila Khalatbari, Maria Ryskina, Rita Frieske, Ryan Cotterell, and Zhijing Jin. 2022 · 2022
Later among the works it cites.
Queer in ai
Hetvi Jethwani, Arjun Subramonian, William Agnew, MaryLena Bleile, Sarthak Arora, Maria Ryskina, and Jeffrey Xiong. 2022 · 2022
Later among the works it cites.
Investigating reasons for disagreement in natural language inference
Original
Nan Jiang and Marie-Catherine de Marneffe. 2022 · 2022
Later among the works it cites.
What if ground truth is subjective? personalized deep neural hate speech detection
Kamil Kanclerz, Marcin Gruza, Konrad Karanowski, Julita Bielaniewcz, Piotr Miłkowski, Jan Kocon, and Przemysław Kazienko. 2022 · 2022
Later among the works it cites.
What do nlp researchers believe? results of the nlp community metasurvey
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, , and Samuel R. Bowman. 2022 · 2022
Later among the works it cites.
Extrinsic evaluation of machine translation metrics
Nikita Moghe, Tom Sherborne, Mark Steedman, and Alexandra Birch. 2022 · 2022
Later among the works it cites.
Don’t blame the annotator: Bias already starts in the annotation instructions
Original
Mihir Parmar, Swaroop Mishra, Mor Geva, and Chitta Baral. 2022 · 2022
Later among the works it cites.
The ’problem’ of human label variation: On ground truth in data, modeling and evaluation
Barbara Plank. 2022 · 2022
Later among the works it cites.
Features or spurious artifacts? data-centric baselines for fair and robust hate speech detection
Alan Ramponi and Sara Tonelli. 2022 · 2022
Later among the works it cites.
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection
Original
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022 · 2022
Later among the works it cites.
The tail wagging the dog: Dataset construction biases of social bias benchmarks
Original
Nikil Roashan Selvam, Sunipa Dev, Daniel Khashabi, Tushar Khot, and Kai-Wei Chang. 2022 · 2022
Later among the works it cites.
Quantifying social biases using templates is unreliable
Original
Preethi Seshadri, Pouya Pezeshkpour, and Sameer Singh. 2022 · 2022
Later among the works it cites.
Talking about large language models
Murray Shanahan. 2022 · 2022
Later among the works it cites.
Task ambiguity in humans and language models
Alex Tamkin, Kunal Handa, Ava Shrestha, and Noah D. Goodman. 2022 · 2022
Later among the works it cites.
Understanding factual errors in summarization: Errors, summarizers, datasets, error detectors
Original
Liyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban, Jiacheng Xu, Semih Yahvuz, Wojciech Kryscinski, Justin F. Rousseau, and Greg Durrett. 2022 · 2022
Later among the works it cites.
Predicting is not understanding: Recognizing and addressing underspecification in machine learning
Damien Teney, Maxime Peyrard, and Ehsan Abbasnejad. 2022 · 2022
Later among the works it cites.
What makes a good and useful summary? Incorporating users in automatic summarization research
Maartje Ter Hoeve, Julia Kiseleva, and Maarten Rijke. 2022 · 2022
Later among the works it cites.
Are ground truth labels reproducible? an empirical study
Ka Wong, Praveen Paritosh, and Kurt Bollacker. 2022 · 2022
Later among the works it cites.
Can we fix the scope for coreference? problems and solutions for benchmarks beyond ontonotes
Original
Amir Zeldes. 2022 · 2022
Later among the works it cites.
Extractive is not faithful: An investigation of broad unfaithfulness problems in extractive summarization
Shiyue Zhang, David Wan, and Mohit Bansal. 2022 · 2022
Later among the works it cites.
Deconstructing nlg evaluation: Evaluation practices, assumptions, and their implications
Original
Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daum’e, Kaheer Suleman, and Alexandra Olteanu. 2022 · 2022
Later among the works it cites.
Dissociating language and thought in large language models: a cognitive perspective
Kyle Mahowald, Anna A. Ivanova, Idan Asher Blank, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. 2023 · 2023
Closest in time.