Fetching the paper…
Reading the bibliography…
In contrast to conventional visual question answering, video-grounded dialog necessitates a profound understanding of both dialog history and video content for accurate response generation.
Focal visual-text attention for memex question answering
Junwei Liang, Lu Jiang, Liangliang Cao, Yannis Kalantidis, Li-Jia Li, and Alexander G Hauptmann. 2019 · 1908
Earlier work this paper cites.
Active scene recognition with vision and language. In ICCV . 810–817
Xiaodong Yu, Cornelia Fermüller, Ching Lik Teo, Yezhou Yang, and Yiannis Aloimonos. 2011 · 2011
Earlier work this paper cites.
The cognitive dialogue: A new model for vision implementing common sense reasoning
Yiannis Aloimonos and Cornelia Fermüller. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks. In NeurIPS . 1–9
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding. In ECCV . 510–526
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016 · 2016
Earlier work this paper cites.
Audio visual scene-aware dialog (avsd) track for natural language generation in dstc7. In AAAI Workshop
Huda Alamri, Chiori Hori, Tim K Marks, Dhruv Batra, and Devi Parikh. 2018 · 2018
Earlier work this paper cites.
Attentive moment retrieval in videos. In The 41st international ACM SIGIR conference on research & development in information retrieval . 15–24
Meng Liu, Xiang Wang, Liqiang Nie, Xiangnan He, Baoquan Chen, and Tat-Seng Chua. 2018 · 2018
Earlier work this paper cites.
End-to-end audio visual scene-aware dialog using multimodal attention-based video features. In ICASSP . 2352–2356
Chiori Hori, Huda Alamri, Jue Wang, Gordon Wichern, Takaaki Hori, Anoop Cherian, Tim K Marks, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das, et al · 2019
Earlier work this paper cites.
Entropy-enhanced multimodal attention model for scene-aware dialogue generation. In AAAI Workshop
Kuan-Yen Lin, Chao-Chun Hsu, Yun-Nung Chen, and Lun-Wei Ku. 2019 · 2019
Earlier work this paper cites.
Recursive visual attention in visual dialog. In CVPR . 6679–6688
Yulei Niu, Hanwang Zhang, Manli Zhang, Jianhong Zhang, Zhiwu Lu, and Ji-Rong Wen. 2019 · 2019
Earlier work this paper cites.
Cmu sinbad’s submission for the dstc7 avsd challenge. In AAAI Workshop
Ramon Sanabria, Shruti Palaskar, and Florian Metze. 2019 · 2019
Earlier work this paper cites.
Reactive multi-stage feature fusion for multimodal dialogue modeling. In AAAI Workshop
Yi-Ting Yeh, Tzu-Chuan Lin, Hsiao-Hua Cheng, Yu-Hsuan Deng, Shang-Yu Su, and Yun-Nung Chen. 2019 · 2019
Earlier work this paper cites.
Multi-step joint-modality attention network for scene-aware dialogue system. In AAAI Workshop
Yun-Wei Chu, Kuan-Yen Lin, Chao-Chun Hsu, and Lun-Wei Ku. 2020 · 2020
Earlier work this paper cites.
Iterative context-aware graph inference for visual dialog. In CVPR . 10055–10064
Dan Guo, Hui Wang, Hanwang Zhang, Zheng-Jun Zha, and Meng Wang. 2020 · 2020
Earlier work this paper cites.
Audio visual scene-aware dialog (avsd) track for natural language generation in DSTC8. In AAAI Workshop
Chiori Hori, Anoop Cherian, Takaaki Hori, and Tim K Marks. 2020 · 2020
Earlier work this paper cites.
Multimodal transformer with pointer network for the dstc8 avsd challenge. In AAAI Workshop
Hung Le and Nancy F Chen. 2020 · 2020
Earlier work this paper cites.
Video-grounded dialogues with pretrained generation language models. In ACL . 5842–5848
Hung Le and Steven CH Hoi. 2020 · 2020
Earlier work this paper cites.
BiST: Bi-directional spatio-temporal reasoning for video-grounded dialogues. In EMNLP . 1846–1859
Hung Le, Doyen Sahoo, Nancy F Chen, and Steven CH Hoi. 2020 · 2020
Earlier work this paper cites.
Dstc8-avsd: Multimodal semantic transformer network with retrieval style word generator. In AAAI Workshop
Hwanhee Lee, Seunghyun Yoon, Franck Dernoncourt, Doo Soon Kim, Trung Bui, and Kyomin Jung. 2020 · 2020
Earlier work this paper cites.
Efficient attention mechanism for visual dialog that can handle all the interactions between multiple inputs. In ECCV . 223–240
Van-Quang Nguyen, Masanori Suganuma, and Takayuki Okatani. 2020 · 2020
Earlier work this paper cites.
MASK-RL: Multiagent Video Object Segmentation Framework Through Reinforcement Learning
Giuseppe Vecchio, Simone Palazzo, Daniela Giordano, Francesco Rundo, and Concetto Spampinato. 2020 · 2020
Earlier work this paper cites.
VD-BERT: A unified vision and dialog transformer with BERT. In EMNLP . 3325–3338
Yue Wang, Shafiq Joty, Michael Lyu, Irwin King, Caiming Xiong, and Steven C.H. Hoi. 2020 · 2020
Cited alongside, same era.
Audio visual scene-aware dialog system using dynamic memory networks. In AAAI Workshop
Huiyuan Xie and Ignacio Iacobacci. 2020 · 2020
Cited alongside, same era.
Memory Augmented Deep Recurrent Neural Network for Video Question Answering
Chengxiang Yin, Jian Tang, Zhiyuan Xu, and Yanzhi Wang. 2020 · 2020
Cited alongside, same era.
Dynamic graph representation learning for video dialog via multi-modal shuffled transformers. In AAAI . 1415–1423
Shijie Geng, Peng Gao, Moitreya Chatterjee, Chiori Hori, Jonathan Le Roux, Yongfeng Zhang, Hongsheng Li, and Anoop Cherian. 2021 · 2021
Cited alongside, same era.
Structured co-reference graph attention for video-grounded dialogue. In AAAI . 1789–1797
Junyeong Kim, Sunjae Yoon, Dahyun Kim, and Chang D Yoo. 2021 · 2021
Cited alongside, same era.
DialogMCF: Multimodal context flow for audio visual scene-aware dialog
Zhe Chen, Hongcheng Liu, and Yu Wang. 2023 · 2023
Closest in time.
Learning multi-turn response selection in grounded dialogues with reinforced knowledge and context distillation
Jiazhan Feng, Chongyang Tao, Xueliang Zhao, and Dongyan Zhao. 2023 · 2023
Closest in time.
Otter: A multi-modal model with in-context instruction tuning
Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. 2023b · 2023
Closest in time.
Video-llava: Learning united visual representation by alignment before projection
Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. 2023 · 2023
Closest in time.
Question-Guided Erasing-Based Spatiotemporal Attention Learning for Video Question Answering
Fei Liu, Jing Liu, Richang Hong, and Hanqing Lu. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning reasoning paths over semantic graphs for video-grounded dialogues. In ICLR
Hung Le, Nancy F Chen, and Steven CH Hoi. 2021 · 2021
Cited alongside, same era.
Bridging text and video: A universal multimodal transformer for audio-visual scene-aware dialog
Zekang Li, Zongjia Li, Jinchao Zhang, Yang Feng, and Jie Zhou. 2021b · 2021
Cited alongside, same era.
Modeling text-visual mutual dependency for multi-modal dialog generation
Shuhe Wang, Yuxian Meng, Xiaofei Sun, Fei Wu, Rongbin Ouyang, Rui Yan, Tianwei Zhang, and Jiwei Li. 2021 · 2021
Cited alongside, same era.
SeqDialN: Sequential visual dialog network in joint visual-linguistic representation space. In AAAI Workshop . 8–17
Liu Yang, Fanqi Meng, Xiao Liu, Ming-Kuang Daniel Wu, Vicent Ying, and James Xu. 2021 · 2021
Cited alongside, same era.
Multimodal dialog system: Relational graph-based context-aware question understanding. In ACM MM . 695–703
Haoyu Zhang, Meng Liu, Zan Gao, Xiaoqiang Lei, Yinglong Wang, and Liqiang Nie. 2021 · 2021
Cited alongside, same era.
UTC: A unified transformer with inter-task contrastive learning for visual dialog. In CVPR . 18103–18112
Cheng Chen, Zhenshan Tan, Qingrong Cheng, Xin Jiang, Qun Liu, Yudong Zhu, and Xiaodong Gu. 2022 · 2022
Cited alongside, same era.
Context-aware graph inference with knowledge distillation for visual dialog
Dan Guo, Hui Wang, and Meng Wang. 2022 · 2022
Cited alongside, same era.
Closest in time.
Vstar: A video-grounded dialogue dataset for situated semantic understanding with scene and topic transitions. In ACL . 5036–5048
Yuxuan Wang, Zilong Zheng, Xueliang Zhao, Jinpeng Li, Yueqian Wang, and Dongyan Zhao. 2023 · 2023
Closest in time.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Hang Zhang, Xin Li, and Lidong Bing. 2023b · 2023
Closest in time.
Attribute-guided collaborative learning for partial person re-identification
Haoyu Zhang, Meng Liu, Yuhong Li, Ming Yan, Zan Gao, Xiaojun Chang, and Liqiang Nie. 2023c · 2023
Closest in time.
A static and dynamic attention framework for multi turn dialogue generation
Weinan Zhang, Yiming Cui, Kaiyan Zhang, Yifa Wang, Qingfu Zhu, Lingzhi Li, and Ting Liu. 2023a · 2023
Closest in time.
AudioVisual Video Summarization
Bin Zhao, Maoguo Gong, and Xuelong Li. 2023b · 2023
Closest in time.
A Weighted Heterogeneous Graph-Based Dialog System
Xinyan Zhao, Liangwei Chen, and Huanhuan Chen. 2023a · 2023
Closest in time.
Multimodal dialog systems with dual knowledge-enhanced generative pretrained language model
Xiaolin Chen, Xuemeng Song, Liqiang Jing, Shuo Li, Linmei Hu, and Liqiang Nie. 2024 · 2024
Closest in time.
Voice-Face Homogeneity Tells Deepfake
Harry Cheng, Yangyang Guo, Tianyi Wang, Qi Li, Xiaojun Chang, and Liqiang Nie. 2024a · 2024
Closest in time.
Towards efficient coarse-grained dialogue response selection
Tian Lan, Xian-Ling Mao, Wei Wei, Xiaoyan Gao, and Heyan Huang. 2024 · 2024
Closest in time.
Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
Zaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen, Dongmei Jiang, and Liqiang Nie. 2024 · 2024
Closest in time.
DKDM: Data-Free Knowledge Distillation for Diffusion Models with Any Architecture
Qianlong Xiang, Miao Zhang, Yuzhang Shang, Jianlong Wu, Yan Yan, and Liqiang Nie. 2024 · 2024
Closest in time.
HCQA@ Ego4D EgoSchema Challenge 2024
Haoyu Zhang, Yuquan Xie, Yisen Feng, Zaijing Li, Meng Liu, and Liqiang Nie. 2024b · 2024
Closest in time.
Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policy. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE
Zaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen, Dongmei Jiang, and Liqiang Nie. 2025 · 2025
Closest in time.
Factor graph attention. In CVPR . 2039–2048
Idan Schwartz, Seunghak Yu, Tamir Hazan, and Alexander G Schwing. 2019 · 2048
Closest in time.
Adversarial vqa: A new benchmark for evaluating the robustness of vqa models. In ICCV . 2042–2051
Linjie Li, Jie Lei, Zhe Gan, and Jingjing Liu. 2021a · 2051
Closest in time.