Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated remarkable capabilities across a broad spectrum of tasks.
Equalizing gender biases in neural machine translation with word embeddings techniques
Joel Escudé Font and Marta R. Costa-jussà · 1901
Earlier work this paper cites.
The second conversational intelligence challenge (convai2)
Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander H. Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, Shrimai Prabhumoye, Alan W. Black, Alexander I. Rudnicky, Jason D. Williams, Joelle Pineau, Mikhail S. Burtsev, and Jason Weston · 1902
Earlier work this paper cites.
Racial bias in hate speech and abusive language detection datasets
Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber · 1905
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Semeval-2016 task 4: Sentiment analysis in twitter
Preslav Nakov, Alan Ritter, Sara Rosenthal, Fabrizio Sebastiani, and Veselin Stoyanov · 1912
Earlier work this paper cites.
Representations of commonsense knowledge
Ernest Davis · 1990
Earlier work this paper cites.
The child’s theory of mind
Henry M Wellman · 1992
Earlier work this paper cites.
Message understanding conference- 6: A brief history
Ralph Grishman and Beth Sundheim · 1996
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme · 2002
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder · 2003
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2003
Earlier work this paper cites.
Common Morality: Deciding What to Do
Bernard Gert · 2004
Earlier work this paper cites.
Core mechanisms in ‘theory of mind’
Alan M Leslie, Ori Friedman, and Tim P German · 2004
Earlier work this paper cites.
Conceptnet—a practical commonsense reasoning tool-kit
Hugo Liu and Push Singh · 2004
Earlier work this paper cites.
Theory of mind
Chris Frith and Uta Frith · 2005
Earlier work this paper cites.
Examining gender and race bias in two hundred sentiment analysis systems
Svetlana Kiritchenko and Saif M. Mohammad · 2005
Earlier work this paper cites.
Liberals and conservatives rely on different sets of moral foundations
Jesse Graham, Jonathan Haidt, and Brian A Nosek · 2009
Earlier work this paper cites.
Impact of news on the commodity market: Dataset and results
Ankur Sinha and Tanmay Khandait · 2009
Earlier work this paper cites.
LEGAL-BERT: the muppets straight out of law school
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos · 2010
Earlier work this paper cites.
Unqovering stereotyping biases via underspecified questions
Tao Li, Tushar Khot, Daniel Khashabi, Ashish Sabharwal, and Vivek Srikumar · 2010
Earlier work this paper cites.
Semeval-2019 task 6: Identifying and categorizing offensive language in social media (offenseval)
Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar · 2010
Earlier work this paper cites.
Isanette: A common and common sense knowledge base for opinion mining
Erik Cambria, Yangqiu Song, Haixun Wang, and Amir Hussain · 2011
Earlier work this paper cites.
The winograd schema challenge
Hector J. Levesque · 2011
Earlier work this paper cites.
Representing general relational knowledge in conceptnet 5
Robyn Speer and Catherine Havasi · 2012
Earlier work this paper cites.
Learning to solve arithmetic word problems with verb categorization
Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman · 2014
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
Pekka Malo, Ankur Sinha, Pekka J. Korhonen, Jyrki Wallenius, and Pyry Takala · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
Panupong Pasupat and Percy Liang · 2015
Earlier work this paper cites.
Solving general arithmetic word problems
Subhro Roy and Dan Roth · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam Tauman Kalai · 2016
Earlier work this paper cites.
MAWPS: A math word problem repository
Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi · 2016
Earlier work this paper cites.
SQuAD: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Overview of the medical question answering task at TREC 2017 liveqa
Asma Ben Abacha, Eugene Agichtein, Yuval Pinter, and Dina Demner-Fushman · 2017
Earlier work this paper cites.
A causal framework for explaining the predictions of black-box sequence-to-sequence models
David Alvarez-Melis and Tommi S. Jaakkola · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
The trouble with bias
Kate Crawford · 2017
Earlier work this paper cites.
https://plato.stanford.edu/archives/sum2017/entries/abduction/ , 2017
Igor Douven · 2017
Earlier work this paper cites.
Search-based neural structured learning for sequential question answering
Mohit Iyyer, Wen-tau Yih, and Ming-Wei Chang · 2017
Earlier work this paper cites.
Going gray, failure to hire, and the ick factor: Analyzing how older bloggers talk about ageism
Amanda Lazar, Mark Diaz, Robin Brewer, Chelsea Kim, and Anne Marie Piper · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji · 2017
Earlier work this paper cites.
Newsqa: A machine comprehension dataset
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman · 2017
Earlier work this paper cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning
Victor Zhong, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the AI2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
XNLI: evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman · 2018
Earlier work this paper cites.
Classification of moral foundations in microblog political discourse
Kristen Johnson and Dan Goldwasser · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth · 2018
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
Tomás Kociský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette · 2018
Earlier work this paper cites.
Www’18 open challenge: Financial opinion mining and question answering
Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? A new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Obtaining reliable human ratings of valence, arousal, and dominance for 20, 000 english words
Saif M. Mohammad · 2018
Earlier work this paper cites.
Reducing gender bias in abusive language detection
Ji Ho Park, Jamin Shin, and Pascale Fung · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018
Earlier work this paper cites.
Mind the GAP: A balanced corpus of gendered ambiguous pronouns
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge · 2018
Earlier work this paper cites.
Constructing datasets for multi-hop reading comprehension across documents
Johannes Welbl, Pontus Stenetorp, and Sebastian Riedel · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Personalizing dialogue agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston · 2018
Earlier work this paper cites.
Benchmarking safe exploration in deep reinforcement learning
Joshua Achiam and Dario Amodei · 2019
Earlier work this paper cites.
Finbert: Financial sentiment analysis with pre-trained language models
Dogu Araci · 2019
Earlier work this paper cites.
Findings of the WMT 2019 biomedical translation shared task: Evaluation for MEDLINE abstracts and biomedical terminologies
Rachel Bawden, Kevin Bretonnel Cohen, Cristian Grozea, Antonio Jimeno-Yepes, Madeleine Kittner, Martin Krallinger, Nancy Mah, Aurélie Névéol, Mariana L. Neves, Felipe Soares, Amy Siu, Karin Verspoor, and Maika Vicente Navarro · 2019
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman · 2019
Earlier work this paper cites.
Finding microaggressions in the wild: A case for locating elusive phenomena in social media posts
Luke Breitfeller, Emily Ahn, David Jurgens, and Yulia Tsvetkov · 2019
Earlier work this paper cites.
https://www.lsac.org/lsat/taking-lsat/test-format/logical-reasoning , 2019
Law School Admission Council · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Addressing age-related bias in sentiment analysis
Mark Díaz, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle · 2019
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston · 2019
Earlier work this paper cites.
Jigsaw unintended bias in toxicity classification
Quan Do · 2019
Earlier work this paper cites.
Ranking generated summaries by correctness: An interesting but challenging application for natural language inference
Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych · 2019
Earlier work this paper cites.
Assessing the factual accuracy of generated text
Ben Goodrich, Vinay Rao, Peter J. Liu, and Mohammad Saleh · 2019
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander H. Miller · 2019
Earlier work this paper cites.
Towards empathetic open-domain conversation models: A new benchmark and dataset
Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau · 2019
Earlier work this paper cites.
Coqa: A conversational question answering challenge
Siva Reddy, Danqi Chen, and Christopher D. Manning · 2019
Earlier work this paper cites.
Enhancing the measurement of social effects by capturing morality
Rezvaneh Rezapour, Saumil H. Shah, and Jana Diesner · 2019
Earlier work this paper cites.
Social iqa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi · 2019
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng · 2019
Earlier work this paper cites.
Evaluating gender bias in machine translation
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Gendered ambiguous pronoun (gap) shared task at the gender bias in nlp workshop 2019
Kellie Webster, Marta R Costa-Jussà, Christian Hardmeier, and Will Radford · 2019
Earlier work this paper cites.
Dialogue natural language inference
Sean Welleck, Jason Weston, Arthur Szlam, and Kyunghyun Cho · 2019
Earlier work this paper cites.
Can neural networks understand monotonicity reasoning?
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos · 2019
Earlier work this paper cites.
HELP: A dataset for identifying shortcomings of neural models in monotonicity reasoning
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos · 2019
Earlier work this paper cites.
Predicting the type and target of offensive posts in social media
Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
"going on a vacation" takes longer than "going for a walk": A study of temporal commonsense understanding
Ben Zhou, Daniel Khashabi, Qiang Ning, and Dan Roth · 2019
Earlier work this paper cites.
PIQA: reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Toward gender-inclusive coreference resolution
Yang Trista Cao and Hal Daumé III · 2020
Earlier work this paper cites.
Hybridqa: A dataset of multi-hop question answering over tabular and textual data
Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang · 2020
Earlier work this paper cites.
On measuring and mitigating biased inferences of word embeddings
Sunipa Dev, Tao Li, Jeff M. Phillips, and Vivek Srikumar · 2020
Earlier work this paper cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Esin Durmus, He He, and Mona T. Diab · 2020
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi · 2020
Earlier work this paper cites.
Towards understanding gender bias in relation extraction
Andrew Gaut, Tony Sun, Shirlyn Tang, Yuxin Huang, Jing Qian, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth M. Belding, Kai-Wei Chang, and William Yang Wang · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Evaluating factuality in generation with dependency-level entailment
Tanya Goyal and Greg Durrett · 2020
Earlier work this paper cites.
A dataset for statutory reasoning in tax law entailment and question answering
Nils Holzenberger, Andrew Blair-Stanek, and Benjamin Van Durme · 2020
Earlier work this paper cites.
Moral foundations twitter corpus: A collection of 35k tweets annotated for moral sentiment
Joe Hoover, Gwenyth Portillo-Wightman, Leigh Yeh, Shreya Havaldar, Aida Mostafazadeh Davani, Ying Lin, Brendan Kennedy, Mohammad Atari, Zahra Kamel, Madelyn Mendlen, et al · 2020
Earlier work this paper cites.
What have we achieved on text summarization?
Dandan Huang, Leyang Cui, Sen Yang, Guangsheng Bao, Kun Wang, Jun Xie, and Yue Zhang · 2020
Earlier work this paper cites.
Social biases in NLP models as barriers for persons with disabilities
Ben Hutchinson, Vinodkumar Prabhakaran, Emily Denton, Kellie Webster, Yu Zhong, and Stephen Denuyl · 2020
Earlier work this paper cites.
Taxinli: Taking a ride up the NLU hill
Pratik Joshi, Somak Aditya, Aalok Sathe, and Monojit Choudhury · 2020
Earlier work this paper cites.
Unifiedqa: Crossing format boundaries with a single QA system
Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi · 2020
Earlier work this paper cites.
QASC: A dataset for question answering via sentence composition
Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Earlier work this paper cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher · 2020
Earlier work this paper cites.
Araweat: Multidimensional analysis of biases in arabic word embeddings
Anne Lauscher, Rafik Takieddin, Simone Paolo Ponzetto, and Goran Glavas · 2020
Earlier work this paper cites.
Does gender matter? towards fairness in dialogue systems
Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang · 2020
Earlier work this paper cites.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang · 2020
Cited alongside, same era.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan T. McDonald · 2020
Cited alongside, same era.
A diverse corpus for evaluating and developing english math word problem solvers
Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su · 2020
Cited alongside, same era.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman · 2020
Cited alongside, same era.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2020
Cited alongside, same era.
OPT: open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer · 2022
Later among the works it cites.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han · 2022
Later among the works it cites.
Towards identifying social bias in dialog systems: Frame, datasets, and benchmarks
Jingyan Zhou, Jiawen Deng, Fei Mi, Yitong Li, Yasheng Wang, Minlie Huang, Xin Jiang, Qun Liu, and Helen Meng · 2022
Later among the works it cites.
The moral integrity corpus: A benchmark for ethical dialogue systems
Caleb Ziems, Jane A. Yu, Yi-Chia Wang, Alon Y. Halevy, and Diyi Yang · 2022
Later among the works it cites.
Frontier AI regulation: Managing emerging risks to public safety
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
Conjnli: Natural language inference over conjunctive sentences
Swarnadeep Saha, Yixin Nie, and Mohit Bansal · 2020
Cited alongside, same era.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi · 2020
Cited alongside, same era.
ALFRED: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2020
Cited alongside, same era.
Can you put it all together: Evaluating conversational agents’ ability to blend skills
Eric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston, and Y-Lan Boureau · 2020
Cited alongside, same era.
Findings of the WMT 2020 shared task on machine translation robustness
Lucia Specia, Zhenhao Li, Juan Miguel Pino, Vishrav Chaudhary, Francisco Guzmán, Graham Neubig, Nadir Durrani, Yonatan Belinkov, Philipp Koehn, Hassan Sajjad, Paul Michel, and Xian Li · 2020
Cited alongside, same era.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis · 2020
Cited alongside, same era.
Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian K. Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager, and Kevin Wolf · 2023
Closest in time.
Evaluating the performance of chatgpt in ophthalmology: An analysis of its successes and shortcomings
Fares Antaki, Samir Touma, Daniel Milad, Jonathan El-Khoury, and Renaud Duval · 2023
Closest in time.
Have llms advanced enough? A challenging problem solving benchmark for large language models
Daman Arora, Himanshu Gaurav Singh, and Mausam · 2023
Closest in time.
Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum
John W. Ayers, Adam Poliak, Mark Dredze, Eric C. Leas, Zechariah Zhu, Jessica B. Kelley, Dennis J. Faix, Aaron M. Goodman, Christopher A. Longhurst, Michael Hogarth, and Davey M. Smith · 2023
Closest in time.
The internal state of an LLM knows when its lying
Amos Azaria and Tom M. Mitchell · 2023
Closest in time.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung · 2023
Closest in time.
Ning Bian, Xianpei Han, Le Sun, Hongyu Lin, Yaojie Lu, and Ben He · 2023
Closest in time.
Can GPT-3 perform statutory reasoning?
Andrew Blair-Stanek, Nils Holzenberger, and Benjamin Van Durme · 2023
Closest in time.
Large language models as tool makers
Tianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen, and Denny Zhou · 2023
Closest in time.
Towards the scalable evaluation of cooperativeness in language models
Alan Chan, Maxime Riché, and Jesse Clifton · 2023
Closest in time.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie · 2023
Closest in time.
Evaluating hallucinations in chinese large language models
Qinyuan Cheng, Tianxiang Sun, Wenwei Zhang, Siyin Wang, Xiangyang Liu, Mozhi Zhang, Junliang He, Mianqiu Huang, Zhangyue Yin, Kai Chen, and Xipeng Qiu · 2023
Closest in time.
I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, and Pengfei Liu · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Closest in time.
Chatgpt goes to law school
Jonathan H Choi, Kristin E Hickman, Amy Monahan, and Daniel Schwarcz · 2023
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel · 2023
Closest in time.
Evaluating language models for mathematics through interactions
Katherine M. Collins, Albert Q. Jiang, Simon Frieder, Lionel Wong, Miri Zilka, Umang Bhatt, Thomas Lukasiewicz, Yuhuai Wu, Joshua B. Tenenbaum, William Hart, Timothy Gowers, Wenda Li, Adrian Weller, and Mateja Jamnik · 2023
Closest in time.
Marta R. Costa-jussà, Pierre Andrews, Eric Smith, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Daniel Licht, and Carleigh Wood · 2023
Closest in time.
Can large language models provide feedback to students? a case study on chatgpt, Apr 2023
Wei Dai, Jionghao Lin, Flora Jin, Tongguang Li, Yi-Shan Tsai, Dragan Gasevic, and Guanliang Chen · 2023
Closest in time.
VNHSGE: vietnamese high school graduation examination dataset for large language models
Xuan-Quy Dao, Ngoc-Bich Le, The-Duy Vo, Xuan-Dung Phan, Bac Bien Ngo, Van-Tien Nguyen, Thi-My-Thanh Nguyen, and Hong Phuoc Nguyen · 2023
Closest in time.
Towards faithful dialogues via focus learning
Yifan Deng, Xingsheng Zhang, Heyan Huang, and Yue Hu · 2023
Closest in time.
How ready are pre-trained abstractive models and llms for legal case judgement summarization?
Aniket Deroy, Kripabandhu Ghosh, and Saptarshi Ghosh · 2023
Closest in time.
Toxicity in chatgpt: Analyzing persona-assigned language models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan · 2023
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Evaluating superhuman models with consistency checks
Lukas Fluri, Daniel Paleka, and Florian Tramèr · 2023
Closest in time.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Closest in time.
Sensitivity and robustness of large language models to prompt in japanese
Chengguang Gan and Tatsunori Mori · 2023
Closest in time.
Understanding social reasoning in language models with language models
Kanishk Gandhi, Jan-Philipp Fränken, Tobias Gerstenberg, and Noah D. Goodman · 2023
Closest in time.
PAL: program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig · 2023
Closest in time.
Trueteacher: Learning factual consistency evaluation with large language models
Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor · 2023
Closest in time.
Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings
Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu · 2023
Closest in time.
An empirical study of metrics to measure representational harms in pre-trained language models
Saghar Hosseini, Hamid Palangi, and Ahmed Hassan Awadallah · 2023
Closest in time.
Geneturing tests gpt models in genomics
Wenpin Hou and Zhicheng Ji · 2023
Closest in time.
Tool documentation enables zero-shot tool-usage with large language models
Cheng-Yu Hsieh, Si-An Chen, Chun-Liang Li, Yasuhisa Fujii, Alexander Ratner, Chen-Yu Lee, Ranjay Krishna, and Tomas Pfister · 2023
Closest in time.
Is chatgpt better than human annotators? potential and limitations of chatgpt in explaining implicit hate speech
Fan Huang, Haewoon Kwak, and Jisun An · 2023
Closest in time.
CBBQ: A chinese bias benchmark dataset curated with human-ai collaboration for large language models
Yufei Huang and Deyi Xiong · 2023
Closest in time.
Multi-dimensional evaluation of text summarization with in-context learning
Sameer Jain, Vaishakh Keshava, Swarnashree Mysore Sathyendra, Patrick Fernandes, Pengfei Liu, Graham Neubig, and Chunting Zhou · 2023
Closest in time.
Yunjie Ji, Yan Gong, Yiping Peng, Chao Ni, Peiyan Sun, Dongyu Pan, Baochang Ma, and Xiangang Li · 2023
Closest in time.
Is chatgpt A good translator? A preliminary study
Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, and Zhaopeng Tu · 2023
Closest in time.
Qiao Jin, Yifan Yang, Qingyu Chen, and Zhiyong Lu · 2023
Closest in time.
Gpt-4 passes the bar exam
Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo · 2023
Closest in time.
On robustness-accuracy characterization of large language models using synthetic datasets
Ching-Yun Ko, Pin-Yu Chen, Payel Das, Yung-Sung Chuang, and Luca Daniel · 2023
Closest in time.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann · 2023
Closest in time.
Giorgi Kokaia, Pratyush Kumar Sinha, Yutong Jiang, and Nozha Boujemaa · 2023
Closest in time.
Comparing code explanations created by students and large language models
Juho Leinonen, Paul Denny, Stephen MacNeil, Sami Sarsa, Seth Bernstein, Joanne Kim, Andrew Tran, and Arto Hellas · 2023
Closest in time.
The diagnostic and triage accuracy of the gpt-3 artificial intelligence model
David M Levine, Rudraksh Tuwani, Benjamin Kompa, Amita Varma, Samuel G. Finlayson, Ateev Mehrotra, and Andrew Beam · 2023
Closest in time.
White-box multi-objective adversarial attack on dialogue generation
Yufei Li, Zexin Li, Yingfan Gao, and Cong Liu · 2023
Closest in time.
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng · 2023
Closest in time.
Agentsims: An open-source sandbox for large language model evaluation
Jiaju Lin, Haoran Zhao, Aochi Zhang, Yiting Wu, Huqiuyue Ping, and Qin Chen · 2023
Closest in time.
Logiqa 2.0 - an improved dataset for logical reasoning in natural language understanding
Hanmeng Liu, Jian Liu, Leyang Cui, Zhiyang Teng, Nan Duan, Ming Zhou, and Yue Zhang · 2023
Closest in time.
Chatgpt as a factual inconsistency evaluator for abstractive text summarization
Zheheng Luo, Qianqian Xie, and Sophia Ananiadou · 2023
Closest in time.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark J. F. Gales · 2023
Closest in time.
Augmented language models: a survey
Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ramakanth Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, Edouard Grave, Yann LeCun, and Thomas Scialom · 2023
Closest in time.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2023
Closest in time.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M. Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel · 2023
Closest in time.
How well do SOTA legal reasoning models support abductive reasoning?
Ha-Thanh Nguyen, Randy Goebel, Francesca Toni, Kostas Stathis, and Ken Satoh · 2023
Closest in time.
Gpt as a financial advisor
Paweł Niszczota and Sami Abbas · 2023
Closest in time.
Capabilities of GPT-4 on medical challenge problems
Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz · 2023
Closest in time.
Chatgpt goes to the operating room: evaluating gpt-4 performance and its potential in surgical education and training in the era of large language models
Namkee Oh, Gyu-Seong Choi, and Woo Yong Lee · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Learning gain differences between chatgpt and human tutor generated algebra hints
Zachary A. Pardos and Shreya Bhandari · 2023
Closest in time.
"why do I feel offended?" - korean dataset for offensive language identification
San-Hee Park, Kang-Min Kim, O-Joun Lee, Youjin Kang, Jaewon Lee, Su-Min Lee, and SangKeun Lee · 2023
Closest in time.
Gorilla: Large language model connected with massive apis
Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez · 2023
Closest in time.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Benjamin Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2023
Closest in time.
Cheng Qian, Chi Han, Yi R. Fung, Yujia Qin, Zhiyuan Liu, and Heng Ji · 2023
Closest in time.
Making language models better tool learners with execution feedback
Shuofei Qiao, Honghao Gui, Huajun Chen, and Ningyu Zhang · 2023
Closest in time.
Webcpm: Interactive web search for chinese long-form question answering
Yujia Qin, Zihan Cai, Dian Jin, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin, Xu Han, Ning Ding, Huadong Wang, Ruobing Xie, Fanchao Qi, Zhiyuan Liu, Maosong Sun, and Jie Zhou · 2023
Closest in time.
Factually consistent summarization via reinforcement learning with textual entailment feedback
Paul Roit, Johan Ferret, Lior Shani, Roee Aharoni, Geoffrey Cideron, Robert Dadashi, Matthieu Geist, Sertan Girgin, Léonard Hussenot, Orgad Keller, Nikola Momchev, Sabela Ramos Garea, Piotr Stanczyk, Nino Vieillard, Olivier Bachem, Gal Elidan, Avinatan Hassidim, Olivier Pietquin, and Idan Szpektor · 2023
Closest in time.
The programmer’s assistant: Conversational interaction with a large language model for software development
Steven I. Ross, Fernando Martinez, Stephanie Houde, Michael J. Muller, and Justin D. Weisz · 2023
Closest in time.
TPTU: task planning and tool usage of large language model-based AI agents
Jingqing Ruan, Yihong Chen, Bin Zhang, Zhiwei Xu, Tianpeng Bao, Guoqing Du, Shiwei Shi, Hangyu Mao, Xingyu Zeng, and Rui Zhao · 2023
Closest in time.
Lost at C: A user study on the security implications of large language model code assistants
Gustavo Sandoval, Hammond Pearce, Teo Nys, Ramesh Karri, Siddharth Garg, and Brendan Dolan-Gavitt · 2023
Closest in time.
Evaluating language-model agents on realistic autonomous tasks
Megan Kinniment Lucas Jun Koba Sato, Haoxing Du, Brian Goodrich, Max Hasin, Lawrence Chan, Luke Harold Miles, Tao R Lin, Hjalmar Wijk, Joel Burget, Aaron Ho, et al · 2023
Closest in time.
Explaining legal concepts with augmented large language models (GPT-4)
Jaromír Savelka, Kevin D. Ashley, Morgan A. Gray, Hannes Westermann, and Huihui Xu · 2023
Closest in time.
Evaluating the moral beliefs encoded in LLMs
Nino Scherrer, Claudia Shi, Amir Feder, and David M. Blei · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Closest in time.
On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning
Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang · 2023
Closest in time.
Prabin Sharma, Kisan Thapa, Dikshya Thapa, Prastab Dhakal, Mala Deep Upadhaya, Santosh Adhikari, and Salik Ram Khanal · 2023
Closest in time.
Model evaluation for extreme risks
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Yoshua Bengio, Paul F. Christiano, and Allan Dafoe · 2023
Closest in time.
Exploring the robustness of large language models for solving programming problems
Atsushi Shirafuji, Yutaka Watanobe, Takumi Ito, Makoto Morishita, Yuki Nakamura, Yusuke Oda, and Jun Suzuki · 2023
Closest in time.
Towards expert-level medical question answering with large language models
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, Mike Schaekermann, Amy Wang, Mohamed Amin, Sami Lachgar, Philip Andrew Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Agüera y Arcas, Nenad Tomasev, Yun Liu, Renee Wong, Christopher Semturs, S. Sara Mahdavi, Joelle K. Barral, Dale R. Webster, Gregory S. Corrado, Yossi Matias, Shekoofeh Azizi, Alan Karthikesalingam, and Vivek Natarajan · 2023
Closest in time.
Restgpt: Connecting large language models with real-world applications via restful apis
Yifan Song, Weimin Xiong, Dawei Zhu, Cheng Li, Ke Wang, Ye Tian, and Sujian Li · 2023
Closest in time.
Robustification of multilingual language models to real-world noise in crosslingual zero-shot settings with robust contrastive pretraining
Asa Cooper Stickland, Sailik Sengupta, Jason Krone, Saab Mansour, and He He · 2023
Closest in time.
A causal framework to quantify the robustness of mathematical reasoning with language models
Alessandro Stolfo, Zhijing Jin, Kumar Shridhar, Bernhard Schölkopf, and Mrinmaya Sachan · 2023
Closest in time.
Recitation-augmented language models
Zhiqing Sun, Xuezhi Wang, Yi Tay, Yiming Yang, and Denny Zhou · 2023
Closest in time.
Evaluating the factual consistency of large language models through news summarization
Derek Tam, Anisha Mascarenhas, Shiyue Zhang, Sarah Kwan, Mohit Bansal, and Colin Raffel · 2023
Closest in time.
Understanding factual errors in summarization: Errors, summarizers, datasets, error detectors
Liyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin F. Rousseau, and Greg Durrett · 2023
Closest in time.
Olid-br: offensive language identification dataset for brazilian portuguese
Douglas Trajano, Rafael H Bordini, and Renata Vieira · 2023
Closest in time.
Plan-and-Solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim · 2023
Closest in time.
Is chatgpt a good teacher coach? measuring zero-shot performance for scoring and providing actionable insights on classroom instruction
Rose E. Wang and Dorottya Demszky · 2023
Closest in time.
Recode: Robustness evaluation of code generation models
Shiqi Wang, Zheng Li, Haifeng Qian, Chenghao Yang, Zijian Wang, Mingyue Shang, Varun Kumar, Samson Tan, Baishakhi Ray, Parminder Bhatia, Ramesh Nallapati, Murali Krishna Ramanathan, Dan Roth, and Bing Xiang · 2023
Closest in time.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Closest in time.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David S. Rosenberg, and Gideon Mann · 2023
Closest in time.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao · 2023
Closest in time.
Do large language models know what they don’t know?
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang · 2023
Closest in time.
KoLA: Carefully benchmarking world knowledge of large language models
Jifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao, Daniel Zhang-li, Xin Lv, Hao Peng, Zijun Yao, Xiaohan Zhang, Hanming Li, Chunyang Li, Zheyuan Zhang, Yushi Bai, Yantao Liu, Amy Xin, Nianyi Lin, Kaifeng Yun, Linlu Gong, Jianhui Chen, Zhili Wu, Yunjia Qi, Weikai Li, Yong Guan, Kaisheng Zeng, Ji Qi, Hailong Jin, Jinxin Liu, Yu Gu, Yuan Yao, Ning Ding, Lei Hou, Zhiyuan Liu, Bin Xu, Jie Tang, and Juanzi Li · 2023
Closest in time.
How well do large language models perform in arithmetic tasks?
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, and Songfang Huang · 2023
Closest in time.
Chatgpt: Unlocking the future of nlp in finance
Adam Zaremba and Ender Demir · 2023
Closest in time.
GLM-130B: an open bilingual pre-trained model
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Zhiyuan Liu, Peng Zhang, Yuxiao Dong, and Jie Tang · 2023
Closest in time.
Measuring massive multitask chinese understanding
Hui Zeng · 2023
Closest in time.
Alignscore: Evaluating factual consistency with A unified alignment function
Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu · 2023
Closest in time.
Xuanyuan 2.0: A large chinese financial chat model with hundreds of billions parameters
Xuanyu Zhang and Qing Yang · 2023
Closest in time.
Robut: A systematic study of table QA robustness against human-annotated adversarial perturbations
Yilun Zhao, Chen Zhao, Linyong Nan, Zhenting Qi, Wenlin Zhang, Xiangru Tang, Boyu Mi, and Dragomir Radev · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Closest in time.
Agieval: A human-centric benchmark for evaluating foundation models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan · 2023
Closest in time.
Webarena: A realistic web environment for building autonomous agents
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig · 2023
Closest in time.
Toolqa: A dataset for LLM question answering with external tools
Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang · 2023
Closest in time.
Evaluation of chatgpt and bert-based models for turkish hate speech detection
Nur Bengisu Çam and Arzucan Özgür · 2023
Closest in time.
Modeling semantic plausibility by injecting world knowledge
Su Wang, Greg Durrett, and Katrin Erk · 2049
Closest in time.
The social impact of natural language processing
Dirk Hovy and Shannon L. Spruit · 2096
Closest in time.