Fetching the paper…
Reading the bibliography…
Given an input video, its associated audio, and a brief caption, the audio-visual scene aware dialog (AVSD) task requires an agent to indulge in a question-answer dialog with a human about the audio-visual content.
Ghosh, S.; Burachas, G.; Ray, A.; and Ziskind, A. 2019 · 1902
Earlier work this paper cites.
MMDetection: Open MMLab Detection Toolbox and Benchmark
Chen, K.; Wang, J.; Pang, J.; Cao, Y.; Xiong, Y.; Li, X.; Sun, S.; Feng, W.; Liu, Z.; Xu, J.; Zhang, Z.; Cheng, D.; Zhu, C.; Cheng, T.; Zhao, Q.; Li, B.; Lu, X.; Zhu, R.; Wu, Y.; Dai, J.; Wang, J.; Shi, J.; Ouyang, W.; Loy, C. C.; and Lin, D. 2019 · 1906
Earlier work this paper cites.
2nd Place Solution to the GQA Challenge 2019
Geng, S.; Zhang, J.; Zhang, H.; Elgammal, A.; and Metaxas, D. N. 2019 · 1907
Earlier work this paper cites.
Reactive multi-stage feature fusion for multimodal dialogue modeling
Yeh, Y.-T.; Lin, T.-C.; Cheng, H.-H.; Deng, Y.-H.; Su, S.-Y.; and Chen, Y.-N. 2019 · 1908
Earlier work this paper cites.
Clotho: An Audio Captioning Dataset
Drossos, K.; Lipping, S.; and Virtanen, T. 2019 · 1910
Earlier work this paper cites.
Deep Audio-Visual Learning: A Survey
Zhu, H.; Luo, M.; Wang, R.; Zheng, A.; and He, R. 2020 · 2001
Earlier work this paper cites.
VideoQA: question answering on news video
Yang, H.; Chaisorn, L.; Zhao, Y.; Neo, S.-Y.; and Chua, T.-S. 2003 · 2003
Earlier work this paper cites.
Character Matters: Video Story Understanding with Character-Aware Relations
Geng, S.; Zhang, J.; Fu, Z.; Gao, P.; Zhang, H.; and de Melo, G. 2020 · 2005
Earlier work this paper cites.
Contrastive Visual-Linguistic Pretraining
Shi, L.; Shuang, K.; Geng, S.; Su, P.; Jiang, Z.; Gao, P.; Fu, Z.; de Melo, G.; and Su, S. 2020b · 2007
Earlier work this paper cites.
GloVe: Global vectors for word representation
Pennington, J.; Socher, R.; and Manning, C. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Chen, X.; Fang, H.; Lin, T.-Y.; Vedantam, R.; Gupta, S.; Dollár, P.; and Zitnick, C. L. 2015 · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
Johnson, J.; Krishna, R.; Stark, M.; Li, L.-J.; Shamma, D.; Bernstein, M.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Sequence to sequence-video to text
Venugopalan, S.; Rohrbach, M.; Donahue, J.; Mooney, R.; Darrell, T.; and Saenko, K. 2015 · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K.; Ba, J.; Kiros, R.; Cho, K.; Courville, A.; Salakhudinov, R.; Zemel, R.; and Bengio, Y. 2015 · 2015
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Fukui, A.; Park, D. H.; Yang, D.; Rohrbach, A.; Darrell, T.; and Rohrbach, M. 2016 · 2016
Earlier work this paper cites.
Structural-rnn: Deep learning on spatio-temporal graphs
Jain, A.; Zamir, A. R.; Savarese, S.; and Saxena, A. 2016 · 2016
Earlier work this paper cites.
Visual Relationship Detection with Language Priors
Lu, C.; Krishna, R.; Bernstein, M.; and Fei-Fei, L. 2016 · 2016
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Cited alongside, same era.
Visual dialog
Das, A.; Kottur, S.; Gupta, K.; Singh, A.; Yadav, D.; Moura, J. M.; Parikh, D.; and Batra, D. 2017 · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Cited alongside, same era.
CNN architectures for large-scale audio classification
Hershey, S.; Chaudhuri, S.; Ellis, D. P.; Gemmeke, J. F.; Jansen, A.; Moore, R. C.; Plakal, M.; Platt, D.; Saurous, R. A.; Seybold, B.; et al. 2017 · 2017
Cited alongside, same era.
Attention-based multimodal fusion for video description
Hori, C.; Hori, T.; Lee, T.-Y.; Zhang, Z.; Harsham, B.; Hershey, J. R.; Marks, T. K.; and Sumi, K. 2017 · 2017
Cited alongside, same era.
TGIF-QA: Toward spatio-temporal reasoning in visual question answering
Multi-step reasoning via recurrent dual attention for visual dialog
Gan, Z.; Cheng, Y.; Kholy, A. E.; Li, L.; Liu, J.; and Gao, J. 2019 · 2019
Later among the works it cites.
Video action transformer network
Girdhar, R.; Carreira, J.; Doersch, C.; and Zisserman, A. 2019 · 2019
Later among the works it cites.
Spatio-temporal action graph networks
Herzig, R.; Levi, E.; Xu, H.; Gao, H.; Brosh, E.; Wang, X.; Globerson, A.; and Darrell, T. 2019 · 2019
Later among the works it cites.
End-to-end audio visual scene-aware dialog using multimodal attention-based video features
Hori, C.; Alamri, H.; Wang, J.; Wichern, G.; Hori, T.; Cherian, A.; Marks, T. K.; Cartillier, V.; Lopes, R. G.; Das, A.; et al. 2019 · 2019
Later among the works it cites.
Multimodal Transformer Networks for End-to-End Video-Grounded Dialogue Systems
Le, H.; Sahoo, D.; Chen, N. F.; and Hoi, S. C. 2019 · 2019
Later among the works it cites.
Self-Attention Graph Pooling
Lee, J.; Lee, I.; and Kang, J. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jang, Y.; Song, Y.; Yu, Y.; Kim, Y.; and Kim, G. 2017 · 2017
Cited alongside, same era.
Visual Genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Cited alongside, same era.
Multimodal Attention for Fusion of Audio and Spatiotemporal Features for Video Description
Hori, C.; Hori, T.; Wichern, G.; Wang, J.; Lee, T.-y.; Cherian, A.; and Marks, T. K. 2018 · 2018
Cited alongside, same era.
Learning conditioned graph structures for interpretable visual question answering
Norcliffe-Brown, W.; Vafeias, S.; and Parisot, S. 2018 · 2018
Cited alongside, same era.
Graph attention networks
Veličković, P.; Cucurull, G.; Casanova, A.; Romero, A.; Lio, P.; and Bengio, Y. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Know more say less: Image captioning based on scene graphs
Li, X.; and Jiang, S. 2019 · 2019
Later among the works it cites.
When does label smoothing help?
Müller, R.; Kornblith, S.; and Hinton, G. E. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Later among the works it cites.
A Simple Baseline for Audio-Visual Scene-Aware Dialog
Schwartz, I.; Schwing, A. G.; and Hazan, T. 2019 · 2019
Later among the works it cites.
Factor Graph Attention
Schwartz, I.; Yu, S.; Hazan, T.; and Schwing, A. G. 2019 · 2019
Later among the works it cites.
Improving grounded natural language understanding through human-robot dialog
Thomason, J.; Padmakumar, A.; Sinapov, J.; Walker, N.; Jiang, Y.; Yedidsion, H.; Hart, J.; Stone, P.; and Mooney, R. J. 2019 · 2019
Later among the works it cites.
Video relationship reasoning using gated spatio-temporal energy graph
Tsai, Y.-H. H.; Divvala, S.; Morency, L.-P.; Salakhutdinov, R.; and Farhadi, A. 2019 · 2019
Later among the works it cites.
Dynamic graph cnn for learning on point clouds
Wang, Y.; Sun, Y.; Liu, Z.; Sarma, S. E.; Bronstein, M. M.; and Solomon, J. M. 2019 · 2019
Later among the works it cites.
Auto-encoding scene graphs for image captioning
Yang, X.; Tang, K.; Zhang, H.; and Cai, J. 2019 · 2019
Later among the works it cites.
Graphical Contrastive Losses for Scene Graph Generation
Zhang, J.; Shih, K. J.; Elgammal, A.; Tao, A.; and Catanzaro, B. 2019 · 2019
Later among the works it cites.
Action Genome: Actions as Composition of Spatio-temporal Scene Graphs
Ji, J.; Krishna, R.; Fei-Fei, L.; and Niebles, J. C. 2020 · 2020
Closest in time.
Video-Grounded Dialogues with Pretrained Generation Language Models
Le, H.; and Hoi, S. C. 2020 · 2020
Closest in time.
Spatio-Temporal Graph for Video Captioning with Knowledge Distillation
Pan, B.; Cai, H.; Huang, D.-A.; Lee, K.-H.; Gaidon, A.; Adeli, E.; and Niebles, J. C. 2020 · 2020
Closest in time.