Fetching the paper…
Reading the bibliography…
The rapid development in the field of Large Language Models (LLMs) has led to a surge in applications that facilitate collaboration among multiple agents to assist humans in their daily tasks.
What makes a good conversation? how controllable attributes affect human judgments
Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. 2019 · 1902
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 1908
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W Cohen, and Xinghua Lu. 2019 · 1909
Earlier work this paper cites.
ALFRED: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. 2019 · 1912
Earlier work this paper cites.
Understanding natural language
Terry Winograd. 1972 · 1972
Earlier work this paper cites.
Optimizing interactive systems via data-driven objectives
Ziming Li, Julia Kiseleva, Alekh Agarwal, Maarten de Rijke, and Ryen W White. 2020 · 2006
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2020b · 2010
Earlier work this paper cites.
Modelling and detecting changes in user satisfaction
Julia Kiseleva, Eric Crestan, Riccardo Brigo, and Roland Dittel. 2014 · 2014
Earlier work this paper cites.
Understanding user satisfaction with intelligent assistants
Julia Kiseleva, Kyle Williams, Jiepu Jiang, Ahmed Hassan Awadallah, Aidan C Crook, Imed Zitouni, and Tasos Anastasakos. 2016b · 2016
Earlier work this paper cites.
Evaluating personal assistants on mobile devices
Julia Kiseleva and Maarten de Rijke. 2017 · 2017
Earlier work this paper cites.
Does that mean you’re happy? rnn-based modeling of user interaction sequences to detect good abandonment
Kyle Williams and Imed Zitouni. 2017 · 2017
Earlier work this paper cites.
Measuring the utility of search engine result pages: an information foraging based measure
Leif Azzopardi, Paul Thomas, and Nick Craswell. 2018 · 2018
Earlier work this paper cites.
Textworld: A learning environment for text-based games
Marc-Alexandre Côté, Akos Kádár, Xingdi Yuan, Ben Kybartas, Tavian Barnes, Emery Fine, James Moore, Matthew Hausknecht, Layla El Asri, Mahmoud Adada, et al. 2019 · 2018
Earlier work this paper cites.
Interactive grounded language understanding in a collaborative environment: Iglu 2021
Julia Kiseleva, Ziming Li, Mohammad Aliannejadi, Shrestha Mohanty, Maartje ter Hoeve, Mikhail Burtsev, Alexey Skrynnik, Artem Zholus, Aleksandr Panov, Kavya Srinet, Arthur Szlam, Yuxuan Sun, Katja Hofmann, Marc-Alexandre Côté, Ahmed Awadallah, Linar Abdrazakov, Igor Churin, Putra Manggala, Kata Naszadi, Michiel van der Meer, and Taewoon Kim. 2022a · 2021
Earlier work this paper cites.
Deus: A data-driven approach to estimate user satisfaction in multi-turn dialogues
Ziming Li, Dookun Park, Julia Kiseleva, Young-Bum Kim, and Sungjin Lee. 2021 · 2021
Cited alongside, same era.
Supporting complex information-seeking tasks with implicit constraints
Ali Ahmadvand, Negar Arabzadeh, Julia Kiseleva, Patricio Figueroa Sanz, Xin Deng, Sujay Jauhar, Michael Gamon, Eugene Agichtein, Ned Friend, et al. 2022 · 2022
Cited alongside, same era.
Interactive grounded language understanding in a collaborative environment: Retrospective on iglu 2022 competition
Julia Kiseleva, Alexey Skrynnik, Artem Zholus, Shrestha Mohanty, Negar Arabzadeh, Marc-Alexandre Côté, Mohammad Aliannejadi, Milagro Teruel, Ziming Li, Mikhail Burtsev, Maartje ter Hoeve, Zoya Volovikova, Aleksandr Panov, Yuxuan Sun, Kavya Srinet, Arthur Szlam, Ahmed Awadallah, Seungeun Rho, Taehwan Kwon, Daniel Wontae Nam, Felipe Bivort Haiek, Edwin Zhang, Linar Abdrazakov, Guo Qingyam, Jason Zhang, and Zhibin Guo. 2022b · 2022
Cited alongside, same era.
Semantic-oriented unlabeled priming for large-scale language models
Yanchen Liu, Timo Schick, and Hinrich Schütze. 2022 · 2022
Evaluating large language models: A comprehensive survey
Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, Deyi Xiong, et al. 2023 · 2023
Later among the works it cites.
Metagpt: Meta programming for multi-agent collaborative framework
Sirui Hong, Xiawu Zheng, Jonathan Chen, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, et al. 2023 · 2023
Later among the works it cites.
Multi-dimensional evaluation of text summarization with in-context learning
Sameer Jain, Vaishakh Keshava, Swarnashree Mysore Sathyendra, Patrick Fernandes, Pengfei Liu, Graham Neubig, and Chunting Zhou. 2023 · 2023
Later among the works it cites.
Alex Liu and Min Sun. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
It is not easy to detect paraphrases: Analysing semantic similarity with antonyms and negation using the new semantoneg benchmark
Teemu Vahtola, Mathias Creutz, and Jörg Tiedemann. 2022 · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022 · 2022
Cited alongside, same era.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023 · 2023
Cited alongside, same era.
Exploring qualitative research using llms
Muneera Bano, Didar Zowghi, and Jon Whittle. 2023 · 2023
Cited alongside, same era.
Ning Bian, Xianpei Han, Le Sun, Hongyu Lin, Yaojie Lu, and Ben He. 2023 · 2023
Cited alongside, same era.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2023 · 2023
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Cited alongside, same era.
Aligning offline metrics and human judgments of value for code generation models
Victor Dibia, Adam Fourney, Gagan Bansal, Forough Poursabzi-Sangdeh, Han Liu, and Saleema Amershi. 2023 · 2023
Cited alongside, same era.
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. 2023 · 2023
Later among the works it cites.
Gaia: a benchmark for general ai assistants
Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
Goal representations for instruction following: A semi-supervised language interface to control
Vivek Myers, Andre Wang He, Kuan Fang, Homer Rich Walke, Philippe Hansen-Estruch, Ching-An Cheng, Mihai Jalobeanu, Andrey Kolobov, Anca Dragan, and Sergey Levine. 2023 · 2023
Later among the works it cites.
Multi-agent collaboration: Harnessing the power of intelligent llm agents
Yashar Talebirad and Amirhossein Nadiri. 2023 · 2023
Later among the works it cites.
Do llms exhibit human-like response biases? a case study in survey design
Lindia Tjuatja, Valerie Chen, Sherry Tongshuang Wu, Ameet Talwalkar, and Graham Neubig. 2023 · 2023
Later among the works it cites.
On the robustness of chatGPT: An adversarial and out-of-distribution perspective
Jindong Wang, Xixu HU, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Wei Ye, Haojun Huang, Xiubo Geng, Binxing Jiao, Yue Zhang, and Xing Xie. 2023 · 2023
Later among the works it cites.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang. 2023 · 2023
Later among the works it cites.
Creative agents: Empowering agents with imagination for creative tasks
Chi Zhang, Penglin Cai, Yuhui Fu, Haoqi Yuan, and Zongqing Lu. 2023 · 2023
Later among the works it cites.
Through the lens of core competency: Survey on evaluation of large language models
Zhuang Ziyu, Chen Qiguang, Ma Longxuan, Li Mingda, Han Yi, Qian Yushan, Bai Haopeng, Zhang Weinan, and Ting Liu. 2023 · 2023
Later among the works it cites.