Fetching the paper…
Reading the bibliography…
We study interactive learning of LLM-based language agents based on user edits made to the agent's output.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I. Levenshtein · 1965
Earlier work this paper cites.
Daniel James Kershaw and R. Koeling · 2008
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
On explore-then-commit strategies
Aurélien Garivier, Tor Lattimore, and Emilie Kaufmann · 2016
Earlier work this paper cites.
Dialogue learning with human-in-the-loop
Jiwei Li, Alexander H. Miller, Sumit Chopra, Marc’Aurelio Ranzato, and Jason Weston · 2016
Earlier work this paper cites.
Dialog-based language learning
Jason Weston · 2016
Earlier work this paper cites.
Generating sentences by editing prototypes
Kelvin Guu, Tatsunori B. Hashimoto, Yonatan Oren, and Percy Liang · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning · 2017
Earlier work this paper cites.
Learning to split and rephrase from wikipedia edit history
Jan A. Botha, Manaal Faruqui, John Alex, Jason Baldridge, and Dipanjan Das · 2018
Earlier work this paper cites.
Sequencer: Sequence-to-sequence learning for end-to-end program repair
Zimin Chen, Steve Kommrusch, Michele Tufano, Louis-Noël Pouchet, Denys Poshyvanyk, and Monperrus Martin · 2018
Earlier work this paper cites.
Pengcheng Yin, Graham Neubig, Miltiadis Allamanis, Marc Brockschmidt, and Alexander L. Gaunt · 2018
Earlier work this paper cites.
On the use of arxiv as a dataset, 2019
Colin B. Clement, Matthew Bierbaum, Kevin P. O’Keeffe, and Alexander A. Alemi · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Learning from dialogue after deployment: Feed yourself, chatbot!
Braden Hancock, Antoine Bordes, Pierre-Emmanuel Mazaré, and Jason Weston · 2019
Earlier work this paper cites.
Argument mining for understanding peer reviews
Xinyu Hua, Mitko Nikolov, Nikhil Badugu, and Lu Wang · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeff Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
A structural model for contextual code changes
Shaked Brody, Uri Alon, and Eran Yahav · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Factual error correction for abstractive summarization models
Mengyao Cao, Yue Dong, Jiapeng Wu, and Jackie Chi Kit Cheung · 2020
Earlier work this paper cites.
Felix Faltings, Michel Galley, Gerold Hintz, Chris Brockett, Chris Quirk, Jianfeng Gao, and Bill Dolan · 2020
Earlier work this paper cites.
Felix: Flexible text editing through tagging and insertion
Jonathan Mallinson, Aliaksei Severyn, Eric Malmi, and Guillermo Garrido · 2020
Earlier work this paper cites.
Variational inference for learning representations of natural language edits
Edison Marrese-Taylor, Machel Reid, and Yutaka Matsuo · 2020
Earlier work this paper cites.
Mpnet: Masked and permuted pre-training for language understanding
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu · 2020
Earlier work this paper cites.
Seq2edits: Sequence transduction using span-level edit operations
Felix Stahlberg and Shankar Kumar · 2020
Earlier work this paper cites.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan J. Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi · 2020
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, S. Arun Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman · 2021
Cited alongside, same era.
Interactive learning from activity description
Khanh X Nguyen, Dipendra Misra, Robert Schapire, Miroslav Dudík, and Patrick Shafto · 2021
Cited alongside, same era.
Learning structural edits via incremental tree transformations
Ziyu Yao, Frank F. Xu, Pengcheng Yin, Huan Sun, and Graham Neubig · 2021
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Generative ai at work
Erik Brynjolfsson, Danielle Li, and Lindsey R Raymond · 2023
Later among the works it cites.
Learning to generate better than your llm
Jonathan D Chang, Kiante Brantley, Rajkumar Ramamurthy, Dipendra Misra, and Wen Sun · 2023
Later among the works it cites.
Llf-bench: Benchmark for interactive learning from language feedback
Ching-An Cheng, Andrey Kolobov, Dipendra Misra, Allen Nie, and Adith Swaminathan · 2023
Later among the works it cites.
Leveraging prefix transfer for multi-intent text revision
Ruining Chong, Cunliang Kong, Liu Wu, Zhenghao Liu, Ziye Jin, Liner Yang, Yange Fan, Hanghang Fan, and Erhong Yang · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An imitation learning curriculum for text editing with non-autoregressive models
Sweta Agrawal and Marine Carpuat · 2022
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback, 2022
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan · 2022
Cited alongside, same era.
Vidhisha Balachandran, Hannaneh Hajishirzi, William Cohen, and Yulia Tsvetkov · 2022
Cited alongside, same era.
Papertweet
Nitsan Bar · 2022
Cited alongside, same era.
Rlprompt: Optimizing discrete text prompts with reinforcement learning
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric P Xing, and Zhiting Hu · 2022
Cited alongside, same era.
Understanding iterative revision from human-written text
Wanyu Du, Vipul Raheja, Dhruv Kumar, Zae Myung Kim, Melissa Lopez, and Dongyeop Kang · 2022
Cited alongside, same era.
Wikimedia downloads
Wikimedia Foundation · 2022
Cited alongside, same era.
Later among the works it cites.
Aries: A corpus of scientific paper edits made in response to peer reviews
Mike D’Arcy, Alexis Ross, Erin Bransom, Bailey Kuehl, Jonathan Bragg, Tom Hope, and Doug Downey · 2023
Later among the works it cites.
Bridging the gap: A survey on integrating (human) feedback for natural language generation
Patrick Fernandes, Aman Madaan, Emmy Liu, António Farinhas, Pedro Henrique Martins, Amanda Bertsch, José G. C. de Souza, Shuyan Zhou, Tongshuang Sherry Wu, Graham Neubig, and André F. T. Martins · 2023
Later among the works it cites.
Swipe: A dataset for document-level simplification of wikipedia pages
Philippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq R. Joty, Caiming Xiong, and Chien-Sheng Wu · 2023
Later among the works it cites.
Automatic prompt rewriting for personalized text generation
Cheng Li, Mingyang Zhang, Qiaozhu Mei, Weize Kong, and Michael Bendersky · 2023
Later among the works it cites.
Second thoughts are best: Learning to re-align with human values from text edits
Ruibo Liu, Chenyan Jia, Ge Zhang, Ziyu Zhuang, Tony X. Liu, and Soroush Vosoughi · 2023
Later among the works it cites.
Pachinko: Patching interpretable qa models through natural language feedback
Chaitanya Malaviya, Subin Lee, Dan Roth, and Mark Yatskar · 2023
Later among the works it cites.
Edit aware representation learning via levenshtein prediction
Edison Marrese-Taylor, Machel Reid, and Alfredo Solano · 2023
Later among the works it cites.
Pearl: Personalizing large language model writing assistants with generation-calibrated retrievers
Sheshera Mysore, Zhuoran Lu, Mengting Wan, Longqi Yang, Steve Menezes, Tina Baghaee, Emmanuel Barajas Gonzalez, Jennifer Neville, and Tara Safavi · 2023
Later among the works it cites.
Learning from free-text human feedback - collect new datasets or extend existing ones?
Dominic Petrak, Nafise Sadat Moosavi, Ye Tian, Nikolai Rozanov, and Iryna Gurevych · 2023
Later among the works it cites.
Diffuser: Diffusion via edit-based reconstruction
Machel Reid, Vincent J. Hellendoorn, and Graham Neubig · 2023
Later among the works it cites.
Training language models with language feedback at scale
J’er’emy Scheurer, Jon Ander Campos, Tomasz Korbak, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez · 2023
Later among the works it cites.
Beyond summarization: Designing ai support for real-world expository writing tasks
Zejiang Shen, Tal August, Pao Siangliulue, Kyle Lo, Jonathan Bragg, Jeff Hammerbacher, Doug Downey, Joseph Chee Chang, and David Sontag · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
Writing with generative ai: Multi-modal and multi-dimensional tools for journalists
Sitong Wang, Lydia B Chilton, and Jeffrey V Nickerson · 2023
Later among the works it cites.
Github copilot is generally available to all developers
Thomas Dohmke · 2024
Closest in time.
Direct language model alignment from online ai feedback
Shangmin Guo, Biao Zhang, Tianlin Liu, Tianqi Liu, Misha Khalman, Felipe Llinares, Alexandre Rame, Thomas Mesnard, Yao Zhao, Bilal Piot, et al · 2024
Closest in time.
A design space for intelligent and interactive writing assistants
Mina Lee, Katy Ilonka Gero, John Joon Young Chung, Simon Buckingham Shum, Vipul Raheja, Hua Shen, Subhashini Venugopalan, Thiemo Wambsganss, David Zhou, Emad A. Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L.C. Guo, Md. Naimul Hoque, Yewon Kim, Seyed Parsa Neshaei, Agnia Sergeyuk, Antonette Shibani, Disha Shrivastava, Lila Shroff, Jessi Stark, S. Sterman, Sitong Wang, Antoine Bosselut, Daniel Buschek, Joseph Chee Chang, Sherol Chen, Max Kreminski, Joonsuk Park, Roy Pea, Eugenia H. Rho, Shannon Zejiang Shen, and Pao Siangliulue · 2024
Closest in time.
Provable interactive learning with hindsight instruction feedback
Dipendra Misra, Aldo Pacchiano, and Robert E Schapire · 2024
Closest in time.