Fetching the paper…
Reading the bibliography…
Many real-world applications of language models (LMs), such as writing assistance and code autocomplete, involve human-LM interaction.
The automatic creation of literature abstracts
Henry P. Luhn · 1958
Earlier work this paper cites.
Response time in man-computer conversational transactions
Robert B Miller · 1968
Earlier work this paper cites.
Evaluation problems in interactive information retrieval
G. Salton · 1970
Earlier work this paper cites.
Metaphors we live by
George Lakoff and Mark Johnson · 1980
Earlier work this paper cites.
Ask for information retrieval: Part I. background and theory
Nicholas J. Belkin, Robert N. Oddy, and Helen M. Brooks · 1982
Earlier work this paper cites.
The information visualizer, an information workspace
Stuart K Card, George G Robertson, and Jock D Mackinlay · 1991
Earlier work this paper cites.
Inside the search process: Information seeking from the user’s perspective
Carol C. Kuhlthau · 1991
Earlier work this paper cites.
Using collaborative filtering to weave an information tapestry
David Goldberg, David Nichols, Brian M Oki, and Douglas Terry · 1992
Earlier work this paper cites.
Information retrieval interaction
Peter Ingwersen · 1992
Earlier work this paper cites.
User interfaces for voice applications
Candace Kamm · 1994
Earlier work this paper cites.
Usability engineering
Jakob Nielsen · 1994
Earlier work this paper cites.
Multiple significance tests: the Bonferroni method
J Martin Bland and Douglas G Altman · 1995
Earlier work this paper cites.
Evaluating Natural Language Processing Systems: An Analysis and Review
Karen Spärck Jones and Julia R. Galliers · 1995
Earlier work this paper cites.
Metaphor and politics
Jeffery Scott Mio · 1997
Earlier work this paper cites.
Challenges for spoken dialogue systems
James Glass · 1999
Earlier work this paper cites.
Patterns of search: Analyzing and modeling web query refinement
Tessa Lau and Eric Horvitz · 1999
Earlier work this paper cites.
Advances in Automatic Text Summarization
Inderjeet Mani · 1999
Earlier work this paper cites.
The TIPSTER SUMMAC text summarization evaluation
Inderjeet Mani, Gary Klein, Lynette Hirschman, Therese Firmin, David House, and Beth Sundheim · 1999
Earlier work this paper cites.
Treebank-3 , 1999
Mitchell Marcus, Beatrice Santorini, Mary Ann Marcinkiewicz, and Ann Taylor · 1999
Earlier work this paper cites.
Automatic summarizing: factors and directions
Karen Spärck Jones · 1999
Earlier work this paper cites.
Understanding figurative language: From metaphor to idioms
Sam Glucksberg and Matthew S McGlone · 2001
Earlier work this paper cites.
A probabilistic approach to solving crossword puzzles
Michael L. Littman, Greg A. Keim, and Noam Shazeer · 2002
Earlier work this paper cites.
Characteristics of question format web queries: an exploratory study
Amanda Spink and H. Cenk Ozmultu · 2002
Earlier work this paper cites.
Voice User Interface Design
Michael Cohen, James P. Giangola, and Jennifer Balogh · 2004
Earlier work this paper cites.
Voice Interaction Design: Crafting the New Conversational Speech Systems
R.A. Harris · 2004
Earlier work this paper cites.
Introduction to recommender systems: Algorithms and evaluation
Joseph A Konstan · 2004
Earlier work this paper cites.
Exploratory search: From finding to understanding
Gary Marchionini · 2006
Earlier work this paper cites.
Proceedings on the workshop on statistical machine translation, 2006
WMT · 2006
Earlier work this paper cites.
Personalized search on the world wide web
Alessandro Micarelli, Fabio Gasparetti, Filippo Sciarrone, and Susan Gauch · 2007
Earlier work this paper cites.
Information re-retrieval: Repeat queries in yahoo’s logs
Jaime Teevan, Eytan Adar, Rosie Jones, and Michael A. S. Potts · 2007
Earlier work this paper cites.
Mining term association patterns from search logs for effective query reformulation
Xuanhai Wang and ChengXiang Zhai · 2008
Earlier work this paper cites.
Analyzing and evaluating query reformulation strategies in web search logs
Jeff Huang and Efthimis N. Efthimiadis · 2009
Earlier work this paper cites.
Patterns of query reformulation during web searching
Bernard J Jansen, Danielle L Booth, and Amanda Spink · 2009
Earlier work this paper cites.
Methods for evaluating interactive information retrieval systems with users
Diane Kelly · 2009
Earlier work this paper cites.
More than cool reason: A field guide to poetic metaphor
George Lakoff and Mark Turner · 2009
Earlier work this paper cites.
Exploratory search: Beyond the query-response paradigm
Ryen W White and Resa A Roth · 2009
Earlier work this paper cites.
Obituary: Fred jelinek
Mark Liberman · 2010
Earlier work this paper cites.
Search in the lost sense of “query”: Question formulation in web search queries and its temporal changes
Bo Pang and Ravi Kumar · 2011
Earlier work this paper cites.
A survey of text summarization techniques
Ani Nenkova and Kathleen McKeown · 2012
Earlier work this paper cites.
Understanding needs embodiment: A theory-guided reanalysis of the role of metaphors and analogies in understanding science
Kai Niebert, Sabine Marsch, and David F Treagust · 2012
Earlier work this paper cites.
Active learning
Burr Settles · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Earlier work this paper cites.
How do users respond to voice input errors? lexical and phonetic query reformulation in voice search
Jiepu Jiang, Wei Jeng, and Daqing He · 2013
Earlier work this paper cites.
How do users grow up along with search engines?: A study of long-term users’ behavior
Jian Liu, Yiqun Liu, Min Zhang, and Shaoping Ma · 2013
Earlier work this paper cites.
The impact of search engine selection and sorting criteria on vaccination beliefs and attitudes: two experiments manipulating google output
Ahmed Allam, Peter Johannes Schulz, Kent Nakamoto, et al · 2014
Earlier work this paper cites.
Dr.fill: Crosswords and an implemented solver for singly weighted csps
Matthew L. Ginsberg · 2014
Earlier work this paper cites.
Questions vs. queries in informational search tasks
Ryen W. White, Matthew Richardson, and Wen tau Tih · 2015
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian V. Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Earlier work this paper cites.
Evaluating visual conversational agents via cooperative human-AI games
Prithvijit Chattopadhyay, Deshraj Yadav, Viraj Prabhu, Arjun Chandrasekaran, Abhishek Das, Stefan Lee, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Socially assistive robotics: Human augmentation versus automation
Maja J Mataric · 2017
Cited alongside, same era.
Task-oriented query reformulation with reinforcement learning
Rodrigo Nogueira and Kyunghyun Cho · 2017
Cited alongside, same era.
Sherlock: A system for interactive summarization of large text collections
PVS Avinesh, Carsten Binnig, Benjamin Hättasch, Christian M Meyer, and Orkan Özyurt · 2018
Cited alongside, same era.
Faithful to the original: Fact-aware neural abstractive summarization
Ziqiang Cao, Furu Wei, Wenjie Li, and Sujian Li · 2018
Cited alongside, same era.
Creative writing with a machine in the loop: Case studies on slogans and stories
Elizabeth Clark, Anne Spencer Ross, Chenhao Tan, Yangfeng Ji, and Noah A. Smith · 2018
Cited alongside, same era.
Challenges in finding metaphorical connections
Katy Gero and Lydia Chilton · 2018
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev · 2021
Later among the works it cites.
AI charades: Language models as interactive game environments
Kevin Frans · 2021
Later among the works it cites.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou · 2021
Later among the works it cites.
Annotating and modeling fine-grained factuality in summarization
Tanya Goyal and Greg Durrett · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi · 2018
Cited alongside, same era.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata · 2018
Cited alongside, same era.
Automated assistance for creative writing with an RNN language model
Melissa Roemmele and Andrew S. Gordon · 2018
Cited alongside, same era.
Toward automated quest generation in text-adventure games
Prithviraj Ammanabrolu, William Broniec, Alex Mueller, Jeremy Paul, and Mark Riedl · 2019
Cited alongside, same era.
Better rewards yield better summaries: Learning to summarise without references
Florian Böhm, Yang Gao, Christian M. Meyer, Ori Shapira, Ido Dagan, and Iryna Gurevych · 2019
Cited alongside, same era.
Gmail smart compose: Real-time assisted writing
Mia Xu Chen, Benjamin N Lee, Gagan Bansal, Yuan Cao, Shuyuan Zhang, Justin Lu, Jackie Tsay, Yinan Wang, Andrew M Dai, Zhifeng Chen, et al · 2019
Cited alongside, same era.
Alexa prize socialbot grand challenge year IV
Dilek Hakkani-Tür · 2021
Later among the works it cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Later among the works it cites.
Further advances in open domain dialog systems in the fourth alexa prize socialbot grand challenge
Shui Hu, Yang Liu, Anna Gottardi, Behnam Hedayatnia, Anju Khatri, Anjali Chadha, Qinlang Chen, Pankaj Rajan, Ali Binici, Varun Somani, Yao Lu, Prerna Dwivedi, Lucy Hu, Hangjie Shi, Sattvik Sahai, Mihail Eric, Karthik Gopalakrishnan, Seokhwan Kim, Spandana Gella, Alexandros Papangelis, Patrick Lange, Di Jin, Nicole Chartier, Mahdi Namazifar, Aishwarya Padmakumar, Sarik Ghazarian, Shereen Oraby, Anjali Narayan-Chen, Yuheng Du, Lauren Stubell, Savanna Stiff, Kate Bland, Arindam Mandal, Reza Ghanadan, and Dilek Hakkani-Tür · 2021
Later among the works it cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams · 2021
Later among the works it cites.
Heuristic evaluation of conversational agents
Raina Langevin, Ross J Lordon, Thi Avrahami, Benjamin R Cowan, Tad Hirsch, and Gary Hsieh · 2021
Later among the works it cites.
Jurassic-1: Technical details and evaluation
Opher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham · 2021
Later among the works it cites.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insignts from training gopher
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving · 2021
Later among the works it cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston · 2021
Later among the works it cites.
Decrypting cryptic crosswords: Semantically complex wordplay
Joshua Rozner, Christopher Potts, and Kyle Mahowald · 2021
Later among the works it cites.
Extending multi-document summarization evaluation to the interactive setting
Ori Shapira, Ramakanth Pasunuru, Hadar Ronen, Mohit Bansal, Yael Amsterdamer, and Ido Dagan · 2021
Later among the works it cites.
Towards interactive language modeling
Maartje ter Hoeve, Evgeny Kharitonov, Dieuwke Hupkes, and Emmanuel Dupoux · 2021
Later among the works it cites.
Commonsense-focused dialogues for response generation: An empirical study
Pei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, and Dilek Hakkani-Tur · 2021
Later among the works it cites.
Measuring progress on scalable oversight for large language models
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan · 2022
Closest in time.
Training language models with language feedback
Jon Ander Campos and Jun Shern · 2022
Closest in time.
Help me write a poem: Instruction tuning as a vehicle for collaborative poetry writing
Tuhin Chakrabarty, Vishakh Padmakumar, and He He · 2022
Closest in time.
PaLM: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, A. Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, B. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, M. Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, S. Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, D. Luan, Hyeontaek Lim, Barret Zoph, A. Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, T. S. Pillai, Marie Pellat, Aitor Lewkowycz, E. Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, K. Meier-Hellstern, D. Eck, J. Dean, Slav Petrov, and Noah Fiedel · 2022
Closest in time.
Aligning offline metrics and human judgments of value of AI-pair programmers
Victor Dibia, Adam Fourney, Gagan Bansal, Forough Poursabzi-Sangdeh, Han Liu, and Saleema Amershi · 2022
Closest in time.
The authenticity gap in human evaluation
Kawin Ethayarajh and Dan Jurafsky · 2022
Closest in time.
The idea machine: LLM-based expansion, rewriting, combination, and suggestion of ideas
Giulia Di Fede, Davide Rocchesso, Steven P. Dow, and Salvatore Andolina · 2022
Closest in time.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, John Kernion, Amanda Askell, Yushi Bai, Saurav Kadavath, Benjamin Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zachary Dodds, T. J. Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom B. Brown, Nicholas Joseph, Sam McCandlish, Christopher Olah, Jared Kaplan, and Jack Clark · 2022
Closest in time.
GEMv2: Multilingual NLG benchmarking in a single line of code
Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran, Alex Wang, Alexandros Papangelis, Aman Madaan, Angelina McMillan-Major, Anna V. Shvets, Ashish Upadhyay, Bingsheng Yao, Bryan Wilie, Chandra Bhagavatula, Chaobin You, Craig Thomson, Cristina Garbacea, Dakuo Wang, Daniel Deutsch, Deyi Xiong, Di Jin, Dimitra Gkatzia, Dragomir Radev, Elizabeth Clark, Esin Durmus, Faisal Ladhak, Filip Ginter, Genta Indra Winata, Hendrik Strobelt, Hiroaki Hayashi, Jekaterina Novikova, Jenna Kanerva, Jenny Chim, Jiawei Zhou, Jordan Clive, Joshua Maynez, João Sedoc, Juraj Juraska, Kaustubh D. Dhole, Khyathi Raghavi Chandu, Leonardo F. R. Ribeiro, Lewis Tunstall, Li Zhang, Mahima Pushkarna, Mathias Creutz, Michael White, Mihir Kale, Moussa Kamal Eddine, Nico Daheim, Nishant Subramani, Ondrej Dusek, Paul Pu Liang, Pawan Sasanka Ammanamanchi, Qinqin Zhu, Ratish Puduppully, Reno Kriz, Rifat Shahriyar, Ronald Cardenas, Saad Mahamood, Salomey Osei, Samuel Cahyawijaya, Sanja vStajner, S’ebastien Montella, Shailza, Shailza Jolly, Simon Mille, Tahmid Hasan, Tianhao Shen, Tosin P. Adewumi, Vikas Raunak, Vipul Raheja, Vitaly Nikolaev, Vivian Tsai, Yacine Jernite, Yi Xu, Yisi Sang, Yixin Liu, and Yufang Hou · 2022
Closest in time.
Alexa, let’s work together: Introducing the first alexa prize taskbot challenge on conversational task assistance
Anna Gottardi, Osman Ipek, Giuseppe Castellucci, Shui Hu, Lavina Vaz, Yao Lu, Anju Khatri, Anjali Chadha, Desheng Zhang, Sattvik Sahai, Prerna Dwivedi, Hangjie Shi, Lucy Hu, Andy Huang, Luke Dai, Bofei Yang, Varun Somani, Pankaj Rajan, Ron Rezac, Michael Johnston, Savanna Stiff, Leslie Ball, David Carmel, Yang Liu, Dilek Hakkani-Tür, Oleg Rokhlenko, Kate Bland, Eugene Agichtein, Reza Ghanadan, and Yoelle Maarek · 2022
Closest in time.
Creative writing with an AI-powered writing assistant: Perspectives from professional writers
Daphne Ippolito, Ann Yuan, Andy Coenen, and Sehmon Burnam · 2022
Closest in time.
Achieving reliable human assessment of open-domain dialogue systems
Tianbo Ji, Yvette Graham, Gareth Jones, Chenyang Lyu, and Qun Liu · 2022
Closest in time.
Handling and presenting harmful text in NLP research
Hannah Rose Kirk, Abeba Birhane, Bertie Vidgen, and Leon Derczynski · 2022
Closest in time.
Faithful or extractive? on mitigating the faithfulness-abstractiveness trade-off in abstractive summarization
Faisal Ladhak, Esin Durmus, He He, Claire Cardie, and Kathleen McKeown · 2022
Closest in time.
CoAuthor: Designing a human-AI collaborative writing dataset for exploring language model capabilities
Mina Lee, Percy Liang, and Qian Yang · 2022
Closest in time.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, D. Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, E. Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel J. Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan S. Kim, Neel Guha, Niladri S. Chatterji, O. Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, S. Ganguli, Tatsunori Hashimoto, Thomas F. Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda · 2022
Closest in time.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp · 2022
Closest in time.
Ambipun: Generating humorous puns with ambiguous context
Anirudh Mittal, Yufei Tian, and Nanyun Peng · 2022
Closest in time.
The brilliance and weirdness of ChatGPT
New York Times · 2022
Closest in time.
Introducing ChatGPT
OpenAI · 2022
Closest in time.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, J. Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, P. Welinder, P. Christiano, J. Leike, and Ryan J. Lowe · 2022
Closest in time.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Closest in time.
Interactive query-assisted summarization via deep reinforcement learning
Ori Shapira, Ramakanth Pasunuru, Mohit Bansal, Ido Dagan, and Yael Amsterdamer · 2022
Closest in time.
BlenderBot 3: a deployed conversational agent that continually learns to responsibly engage
Kurt Shuster, Jing Xu, Mojtaba Komeili, Da Ju, Eric Michael Smith, Stephen Roller, Megan Ung, Moya Chen, Kushal Arora, Joshua Lane, Morteza Behrooz, William Ngan, Spencer Poff, Naman Goyal, Arthur Szlam, Y-Lan Boureau, Melanie Kambadur, and Jason Weston · 2022
Closest in time.
Human evaluation of conversations is an open problem: comparing the sensitivity of various methods for evaluating dialogue agents
Eric Smith, Orion Hsu, Rebecca Qian, Stephen Roller, Y-Lan Boureau, and Jason Weston · 2022
Closest in time.
LaMDA: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le · 2022
Closest in time.
Automated crossword solving
Eric Wallace, Nicholas Tomlin, Albert Xu, Kevin Yang, Eshaan Pathak, Matthew L. Ginsberg, and Dan Klein · 2022
Closest in time.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2022
Closest in time.
Who wrote this? how smart replies impact language and agency in the workplace
Kilian Wenker · 2022
Closest in time.
Missing information, unresponsive authors, experimental flaws: The impossibility of assessing the reproducibility of previous human evaluations in NLP
Anya Belz, Craig Thomson, and Ehud Reiter · 2023
Closest in time.
Choice over control: How users write with large language models using diegetic and non-diegetic prompting
Hai Dang, Sven Goller, Florian Lehmann, and Daniel Buschek · 2023
Closest in time.
AlpacaFarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Co-writing with opinionated language models affects users’ views
Maurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson, and Mor Naaman · 2023
Closest in time.
Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP
Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia · 2023
Closest in time.
G-Eval: NLG evaluation using GPT-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein · 2023
Closest in time.
Don’t blame the annotator: Bias already starts in the annotation instructions
Mihir Parmar, Swaroop Mishra, Mor Geva, and Chitta Baral · 2023
Closest in time.