Fetching the paper…
Reading the bibliography…
In this paper, we propose a Vision-Audio-Language Omni-peRception pretraining model (VALOR) for multi-modal understanding and generation.
J. Cho, J. Lei, H. Tan, and M. Bansal, “Unifying vision-and-language tasks via text generation,” in
1942
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in
2002
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in
2005
Earlier work this paper cites.
S. Banerjee, A. Lavie
2005
Earlier work this paper cites.
D. Chen and W. B. Dolan, “Collecting highly parallel data for paraphrase evaluation,” in
2011
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. Berg, “Im2text: Describing images using 1 million captioned photographs,”
2011
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in
2014
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,”
2015
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in
2015
Earlier work this paper cites.
J. Xu, T. Mei, T. Yao, and Y. Rui, “Msr-vtt: A large video description dataset for bridging video and language,” in
2016
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “Spice: Semantic propositional image caption evaluation,” in
2016
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals
2017
Earlier work this paper cites.
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with natural language,” in
2017
Earlier work this paper cites.
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles, “Dense-captioning events in videos,” in
2017
Earlier work this paper cites.
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele, “Movie description,”
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma
2017
Earlier work this paper cites.
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang, “Video question answering via gradually refined attention over appearance and motion,” in
2017
Earlier work this paper cites.
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim, “Tgif-qa: Toward spatio-temporal reasoning in visual question answering,” in
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in
2017
Earlier work this paper cites.
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in
2017
Earlier work this paper cites.
J.-H. Kim, J. Jun, and B.-T. Zhang, “Bilinear attention networks,”
2018
Earlier work this paper cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in
2018
Earlier work this paper cites.
L. Zhou, C. Xu, and J. J. Corso, “Towards automatic learning of procedures from web instructional videos,” in
2018
Earlier work this paper cites.
Y. Yan, N. Zhuang, B. Ni, J. Zhang, M. Xu, Q. Zhang, Z. Zhang, S. Cheng, Q. Tian, Y. Xu
2019
Earlier work this paper cites.
B. Wang, L. Ma, W. Zhang, W. Jiang, J. Wang, and W. Liu, “Controllable video captioning with pos sequence guidance based on gated fusion network,” in
2019
Earlier work this paper cites.
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,” in
2019
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,”
2019
Earlier work this paper cites.
X. Wang, J. Wu, J. Chen, L. Li, Y.-F. Wang, and W. Y. Wang, “Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,” in
2019
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in
2019
Earlier work this paper cites.
Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao, “Activitynet-qa: A dataset for understanding complex web videos via question answering,” in
2019
Earlier work this paper cites.
L. Peng, Y. Yang, Z. Wang, Z. Huang, and H. T. Shen, “Mra-net: Improving vqa via multi-modal relation attention network,”
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell
2020
Earlier work this paper cites.
A. Miech, J.-B. Alayrac, L. Smaira, I. Laptev, J. Sivic, and A. Zisserman, “End-to-end learning of visual representations from uncurated instructional videos,” in
2020
Earlier work this paper cites.
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “Uniter: Universal image-text representation learning,” in
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
V. Gabeur, C. Sun, K. Alahari, and C. Schmid, “Multi-modal transformer for video retrieval,” in
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
Z. Zhang, Y. Shi, C. Yuan, B. Li, P. Wang, W. Hu, and Z.-J. Zha, “Object relational graph with teacher-recommended learning for video captioning,” in
2020
Cited alongside, same era.
J. Jiang, Z. Chen, H. Lin, X. Zhao, and Y. Gao, “Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering,” in
2020
Cited alongside, same era.
X. Xu, H. Dinkel, M. Wu, and K. Yu, “A crnn-gru based reinforcement learning approach to audio captioning.” in
2020
Cited alongside, same era.
K. Chen, Y. Wu, Z. Wang, X. Zhang, F. Nian, S. Li, and X. Shao, “Audio captioning based on transformer and pre-trained cnn.” in
2020
2022
Later among the works it cites.
2022
Later among the works it cites.
Z. Yang, Z. Gan, J. Wang, X. Hu, F. Ahmed, Z. Liu, Y. Lu, and L. Wang, “Unitab: Unifying text and box outputs for grounded vision-language modeling,” in
2022
Later among the works it cites.
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang, “Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,” in
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
P. He, X. Liu, J. Gao, and W. Chen, “Deberta: Decoding-enhanced bert with disentangled attention,”
2020
Cited alongside, same era.
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei
2020
Cited alongside, same era.
J. Dong, X. Li, C. Xu, X. Yang, G. Yang, X. Wang, and M. Wang, “Dual encoding for video retrieval by text,”
2021
Cited alongside, same era.
F. Liu, X. Wu, C. You, S. Ge, Y. Zou, and X. Sun, “Aligning source visual and target language domains for unpaired video captioning,”
2021
Cited alongside, same era.
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” in
2021
Cited alongside, same era.
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut, “Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,” in
2021
Cited alongside, same era.
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,”
2021
Cited alongside, same era.
2022
Later among the works it cites.
X. Zhu, J. Zhu, H. Li, X. Wu, H. Li, X. Wang, and J. Dai, “Uni-perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Z. Yang, Y. Fang, C. Zhu, R. Pryzant, D. Chen, Y. Shi, Y. Xu, Y. Qian, M. Gao, Y.-L. Chen
2022
Later among the works it cites.
G. Li, Y. Wei, Y. Tian, C. Xu, J.-R. Wen, and D. Hu, “Learning to answer questions in dynamic audio-visual scenarios,” in
2022
Later among the works it cites.
P. Yang, X. Wang, X. Duan, H. Chen, R. Hou, C. Jin, and W. Zhu, “Avqa: A dataset for audio-visual question answering on videos,” in
2022
Later among the works it cites.
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” in
2022
Later among the works it cites.
Y. Gong, C.-I. Lai, Y.-A. Chung, and J. Glass, “Ssast: Self-supervised audio spectrogram transformer,” in
2022
Later among the works it cites.
Y. Ge, Y. Ge, X. Liu, D. Li, Y. Shan, X. Qie, and P. Luo, “Bridging video-text retrieval with multiple choice questions,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Wang, Y. Ge, G. Cai, R. Yan, X. Lin, Y. Shan, X. Qie, and M. Z. Shou, “Object-aware video-language pre-training for retrieval,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Lei, T. L. Berg, and M. Bansal, “Revealing single frame bias for video-and-language learning,”
2022
Later among the works it cites.
V. Gabeur, A. Nagrani, C. Sun, K. Alahari, and C. Schmid, “Masking modalities for cross-modal video retrieval,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
H. Luo, L. Ji, M. Zhong, Y. Chen, W. Lei, N. Duan, and T. Li, “Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning,”
2022
Later among the works it cites.
Y. Liu, P. Xiong, L. Xu, S. Cao, and Q. Jin, “Ts2-net: Token shift and selection transformer for text-video retrieval,” in
2022
Later among the works it cites.
Y. Ma, G. Xu, X. Sun, M. Yan, J. Zhang, and R. Ji, “X-clip: End-to-end multi-grained contrastive learning for video-text retrieval,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Min, W. Kong, R.-C. Tu, D. Gong, C. Cai, W. Zhao, C. Liu, S. Zheng, H. Wang, Z. Li
2022
Later among the works it cites.
Y. Wang, K. Li, Y. Li, Y. He, B. Huang, Z. Zhao, H. Zhang, J. Xu, Y. Liu, Z. Wang
2022
Later among the works it cites.
K. Lin, L. Li, C.-C. Lin, F. Ahmed, Z. Gan, Z. Liu, Y. Lu, and L. Wang, “Swinbert: End-to-end transformers with sparse attention for video captioning,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Koh, X. Fuzhao, and C. E. Siong, “Automated audio captioning using transfer learning and reconstruction latent space similarity regularization,” in
2022
Later among the works it cites.
X. Liu, X. Mei, Q. Huang, J. Sun, J. Zhao, H. Liu, M. D. Plumbley, V. Kilic, and W. Wang, “Leveraging pre-trained bert for audio captioning,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
Z.-Y. Dou, Y. Xu, Z. Gan, J. Wang, S. Wang, L. Wang, C. Zhu, P. Zhang, L. Yuan, N. Peng
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Hu, Z. Gan, J. Wang, Z. Yang, Z. Liu, Y. Lu, and L. Wang, “Scaling up vision-language pre-training for image captioning,” in
2022
Later among the works it cites.
H. Yun, Y. Yu, W. Yang, K. Lee, and G. Kim, “Pano-avqa: Grounded audio-visual question answering on 360deg videos,” in
2041
Closest in time.
L. Li, Y.-C. Chen, Y. Cheng, Z. Gan, L. Yu, and J. Liu, “Hero: Hierarchical encoder for video+ language omni-representation pre-training,” in
2065
Closest in time.