Fetching the paper…
Reading the bibliography…
Recently, substantial advancements in pre-trained vision-language models have greatly enhanced the capabilities of multi-modal dialog systems.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2017
Earlier work this paper cites.
“Multi-modal open-domain dialogue,”
Kurt Shuster, Eric Michael Smith, Da Ju, and Jason Weston, · 2020
Earlier work this paper cites.
“Image-chat: Engaging grounded conversations,”
Kurt Shuster, Samuel Humeau, Antoine Bordes, and Jason Weston, · 2020
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., · 2020
Earlier work this paper cites.
Xiaoxue Zang, Lijuan Liu, Maria Wang, Yang Song, Hao Zhang, and Jindong Chen, · 2021
Earlier work this paper cites.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Earlier work this paper cites.
“The power of scale for parameter-efficient prompt tuning,”
Brian Lester, Rami Al-Rfou, and Noah Constant, · 2021
Earlier work this paper cites.
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang, · 2021
Cited alongside, same era.
“Prefix-tuning: Optimizing continuous prompts for generation,”
Xiang Lisa Li and Percy Liang, · 2021
Cited alongside, same era.
“Multimodal dialogue response generation,”
Qingfeng Sun, Yujing Wang, Can Xu, Kai Zheng, Yaming Yang, Huang Hu, Fei Xu, Jessica Zhang, Xiubo Geng, and Daxin Jiang, · 2022
Cited alongside, same era.
“Flamingo: a visual language model for few-shot learning,”
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al., · 2022
Cited alongside, same era.
“OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,”
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang, · 2022
“Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,”
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi, · 2022
Later among the works it cites.
“Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,”
Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, and Furu Wei, · 2022
Later among the works it cites.
“Scaling instruction-finetuned language models,”
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al., · 2022
Later among the works it cites.
“Mmdialog: A large-scale multi-turn dialogue dataset towards multi-modal open-domain conversation,”
Jiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao, Yaming Yang, Chongyang Tao, Dongyan Zhao, and Qingwei Lin, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Instance-aware prompt learning for language understanding and generation,”
Feihu Jin, Jinliang Lu, Jiajun Zhang, and Chengqing Zong, · 2022
Cited alongside, same era.
“Late prompt tuning: A late prompt could be better than many prompts,”
Xiangyang Liu, Tianxiang Sun, Xuanjing Huang, and Xipeng Qiu, · 2022
Cited alongside, same era.
OpenAI, · 2023
Later among the works it cites.
“Image as a foreign language: Beit pretraining for vision and vision-language tasks,”
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al., · 2023
Later among the works it cites.
“Pace: Unified multi-modal dialogue pre-training with progressive and compositional experts,”
Yunshui Li, Binyuan Hui, ZhiChao Yin, Min Yang, Fei Huang, and Yongbin Li, · 2023
Later among the works it cites.