Fetching the paper…
Reading the bibliography…
Humans share a wide variety of images related to their personal experiences within conversations via instant messaging tools.
Clevr-dialog: A diagnostic dataset for multi-round reasoning in visual dialog
Satwik Kottur, José MF Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach. 2019 · 1903
Earlier work this paper cites.
Dialogcc: An automated pipeline for creating high-quality multi-modal dialogue dataset
Young-Jun Lee, Byungsoo Ko, Han-Gyu Kim, Jonghwan Hyeon, and Ho-Jin Choi. 2024c · 1963
Earlier work this paper cites.
Conversation as a system of social interaction
Rauni Myllyniemi. 1986 · 1986
Earlier work this paper cites.
Four different perspectives on human–computer interaction
John Kammersgaard. 1988 · 1988
Earlier work this paper cites.
To feel or not to feel: The role of affect in human–computer interaction
Eva Hudlicka. 2003 · 2003
Earlier work this paper cites.
Multimodal human–computer interaction: A survey
Alejandro Jaimes and Nicu Sebe. 2007 · 2007
Earlier work this paper cites.
The function of fiction is the abstraction and simulation of social experience
Raymond A Mar and Keith Oatley. 2008 · 2008
Earlier work this paper cites.
Openvidial: A large-scale, open-domain dialogue dataset with visual contexts
Yuxian Meng, Shuhe Wang, Qinghong Han, Xiaofei Sun, Fei Wu, Rui Yan, and Jiwei Li. 2020 · 2012
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
A diagram is worth a dozen images
Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. 2016 · 2016
Earlier work this paper cites.
Photographs as things–photographs of things. a texto-material perspective on photo-sharing practices
Katharina Lobinger. 2016 · 2016
Earlier work this paper cites.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José MF Moura, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Image-grounded conversations: Multimodal context for natural question and response generation
Nasrin Mostafazadeh, Chris Brockett, Bill Dolan, Michel Galley, Jianfeng Gao, Georgios P Spithourakis, and Lucy Vanderwende. 2017 · 2017
Earlier work this paper cites.
Visual reference resolution using attention memory for visual dialog
Paul Hongsuck Seo, Andreas Lehrmann, Bohyung Han, and Leonid Sigal. 2017 · 2017
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Earlier work this paper cites.
Image chat: Engaging grounded conversations
Kurt Shuster, Samuel Humeau, Antoine Bordes, and Jason Weston. 2018 · 2018
Earlier work this paper cites.
Visual semantic reasoning for image-text matching
Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu. 2019 · 2019
Earlier work this paper cites.
(comet-) atomic 2020: On symbolic and neural commonsense knowledge graphs
Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, and Yejin Choi. 2021 · 2020
Earlier work this paper cites.
Just say no: Analyzing the stance of neural dialogue generation in offensive contexts
Ashutosh Baheti, Maarten Sap, Alan Ritter, and Mark Riedl. 2021 · 2021
Earlier work this paper cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. 2021 · 2021
Earlier work this paper cites.
Redcaps: Web-curated image-text data created by the people, for the people
Karan Desai, Gaurav Kaul, Zubin Aysola, and Justin Johnson. 2021 · 2021
Earlier work this paper cites.
Constructing multi-modal dialogue dataset by replacing text with semantically relevant images
Nyoungwoo Lee, Suwon Shin, Jaegul Choo, Ho-Jin Choi, and Sung-Hyun Myaeng. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Cited alongside, same era.
Symbolic knowledge distillation: from general language models to commonsense models
Peter West, Chandra Bhagavatula, Jack Hessel, Jena D Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Beyond goldfish memory: Long-term open-domain conversation
Jing Xu, Arthur Szlam, and Jason Weston. 2021 · 2021
Cited alongside, same era.
Photochat: A human-human dialogue dataset with photo sharing behavior for joint image-text modeling
Fantom: A benchmark for stress-testing machine theory of mind in interactions
Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras, Gunhee Kim, Yejin Choi, and Maarten Sap. 2023 · 2023
Later among the works it cites.
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. 2023 · 2023
Later among the works it cites.
Large language models can share images, too!
Young-Jun Lee, Jonghwan Hyeon, and Ho-Jin Choi. 2023 · 2023
Later among the works it cites.
Photomaker: Customizing realistic human photos via stacked id embedding
Zhen Li, Mingdeng Cao, Xintao Wang, Zhongang Qi, Ming-Ming Cheng, and Ying Shan. 2023 · 2023
Later among the works it cites.
Nlpositionality: Characterizing design biases of datasets and models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiaoxue Zang, Lijuan Liu, Maria Wang, Yang Song, Hao Zhang, and Jindong Chen. 2021 · 2021
Cited alongside, same era.
Mmchat: Multi-modal chat dataset on social media
Yinhe Zheng, Guanyi Chen, Xin Liu, and Ke Lin. 2021 · 2021
Cited alongside, same era.
Mmdialog: A large-scale multi-turn dialogue dataset towards multi-modal open-domain conversation
Jiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao, Yaming Yang, Chongyang Tao, Dongyan Zhao, and Qingwei Lin. 2022 · 2022
Cited alongside, same era.
Seungju Han, Beomsu Kim, Jin Yong Yoo, Seokjun Seo, Sangbum Kim, Enkhbayar Erdenee, and Buru Chang. 2022 · 2022
Cited alongside, same era.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022 · 2022
Cited alongside, same era.
Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities
Mina Lee, Percy Liang, and Qian Yang. 2022a · 2022
Cited alongside, same era.
Instagram photo sharing and its relationships with social connectedness, loneliness, and well-being
Julie Maclean, Yeslam Al-Saggaf, and Rachel Hogg. 2022 · 2022
Cited alongside, same era.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022 · 2022
Cited alongside, same era.
Sebastin Santy, Jenny T Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap. 2023 · 2023
Later among the works it cites.
Introbot: Exploring the use of chatbot-assisted familiarization in online collaborative groups
Donghoon Shin, Soomin Kim, Ruoxi Shang, Joonhwan Lee, and Gary Hsieh. 2023 · 2023
Later among the works it cites.
Cue-cot: Chain-of-thought prompting for responding to in-depth dialogue questions with llms
Hongru Wang, Rui Wang, Fei Mi, Yang Deng, Zezhong Wang, Bin Liang, Ruifeng Xu, and Kam-Fai Wong. 2023 · 2023
Later among the works it cites.
Mind the gap between conversations for improved long-term dialogue generation
Qiang Zhang, Jason Naradowsky, and Yusuke Miyao. 2023 · 2023
Later among the works it cites.
Minigpt-5: Interleaved vision-and-language generation via generative vokens
Kaizhi Zheng, Xuehai He, and Xin Eric Wang. 2023 · 2023
Later among the works it cites.
Sotopia: Interactive evaluation for social intelligence in language agents
Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, et al. 2023 · 2023
Later among the works it cites.
Magid: An automated pipeline for generating synthetic multi-modal datasets
Hossein Aboutalebi, Hwanjun Song, Yusheng Xie, Arshit Gupta, Justin Sun, Hang Su, Igor Shalyminov, Nikolaos Pappas, Siffi Singh, and Saab Mansour. 2024 · 2024
Closest in time.
Llama 3 model card
AI@Meta. 2024 · 2024
Closest in time.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2024 · 2024
Closest in time.
Generating images with multimodal language models
Jing Yu Koh, Daniel Fried, and Russ R Salakhutdinov. 2024 · 2024
Closest in time.
Sdxl-lightning: Progressive adversarial diffusion distillation
Shanchuan Lin, Anran Wang, and Xiao Yang. 2024 · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024 · 2024
Closest in time.
Evaluating very long-term conversational memory of llm agents
Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang. 2024 · 2024
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2024 · 2024
Closest in time.
Measuring multimodal mathematical reasoning with math-vision dataset
Ke Wang, Junting Pan, Weikang Shi, Zimu Lu, Mingjie Zhan, and Hongsheng Li. 2024 · 2024
Closest in time.
Social skill training with large language models
Diyi Yang, Caleb Ziems, William Held, Omar Shaikh, Michael S Bernstein, and John Mitchell. 2024 · 2024
Closest in time.