Fetching the paper…
Reading the bibliography…
A major challenge in text-video and text-audio retrieval is the lack of large-scale training data.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Semantic-audio retrieval
Malcolm Slaney · 2002
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Large-scale content-based audio retrieval from text queries
Gal Chechik, Eugene Ie, Martin Rehn, Samy Bengio, and Dick Lyon · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
Marcus Rohrbach, Sikandar Amin, Mykhaylo Andriluka, and Bernt Schiele · 2012
Earlier work this paper cites.
Freesound technical demo
Frederic Font, Gerard Roma, and Xavier Serra · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
VQA: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Temporal localization of fine-grained actions in videos by domain transfer from web images
Chen Sun, Sanketh Shetty, Rahul Sukthankar, and Ram Nevatia · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Cross modal distillation for supervision transfer
Saurabh Gupta, Judy Hoffman, and Jitendra Malik · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
Show and tell: Lessons learned from the 2015 MSCOCO image captioning challenge
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2016
Earlier work this paper cites.
Learning deep structure-preserving image-text embeddings
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Image captioning with semantic attention
Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Movie description
Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele · 2017
Earlier work this paper cites.
Emotion recognition in speech using cross-modal transfer in the wild
Samuel Albanie, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Cited alongside, same era.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason Corso · 2018
Cited alongside, same era.
Improving the performance of unimodal dynamic hand-gesture recognition with multimodal training
Mahdi Abavisani, Hamid Reza Vaezi Joze, and Vishal M Patel · 2019
Cited alongside, same era.
Multimodal and multi-view models for emotion recognition
Gustavo Aguilar, Viktor Rozgic, Weiran Wang, and Chao Wang · 2019
Cited alongside, same era.
Motion guided spatial attention for video captioning
Shaoxiang Chen and Yu-Gang Jiang · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Later among the works it cites.
Speech2action: Cross-modal supervision for action recognition
Arsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar, Cordelia Schmid, and Andrew Zisserman · 2020
Later among the works it cites.
Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander Hauptmann, João Henriques, and Andrea Vedaldi · 2020
Later among the works it cites.
AVLnet: Learning audio-visual language representations from instructional videos
Andrew Rouditchenko, Angie Boggust, David Harwath, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Rogerio Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Cross modal audio search and retrieval with joint embeddings based on text and audio
Benjamin Elizalde, Shuayb Zarar, and Bhiksha Raj · 2019
Cited alongside, same era.
Large-scale weakly-supervised pre-training for video action recognition
Deepti Ghadiyaram, Du Tran, and Dhruv Mahajan · 2019
Cited alongside, same era.
Distinit: Learning video representations without a single labeled video
Rohit Girdhar, Du Tran, Lorenzo Torresani, and Deva Ramanan · 2019
Cited alongside, same era.
Joint syntax representation learning and visual cue translation for video captioning
Jingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo, and Yunde Jia · 2019
Cited alongside, same era.
Graph-rise: Graph-regularized image semantic embedding
Da-Cheng Juan, Chun-Ta Lu, Zhen Li, Futang Peng, Aleksei Timofeev, Yi-Ting Chen, Yaxi Gao, Tom Duerig, Andrew Tomkins, and Sujith Ravi · 2019
Cited alongside, same era.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim · 2019
Cited alongside, same era.
Jonathan C Stroud, Zhichao Lu, Chen Sun, Jia Deng, Rahul Sukthankar, Cordelia Schmid, and David A Ross · 2020
Later among the works it cites.
Towards fairer datasets: Filtering and balancing the distribution of the people subtree in the imagenet hierarchy
Kaiyu Yang, Klint Qinami, Li Fei-Fei, Jia Deng, and Olga Russakovsky · 2020
Later among the works it cites.
Object relational graph with teacher-recommended learning for video captioning
Ziqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li, Peijin Wang, Weiming Hu, and Zheng-Jun Zha · 2020
Later among the works it cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Later among the works it cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe · 2021
Later among the works it cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut · 2021
Later among the works it cites.
Improving video-text retrieval by multi-stream corpus alignment and dual softmax loss, 2021
Xing Cheng, Hezheng Lin, Xiangyu Wu, Fan Yang, and Dong Shen · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Clip2video: Mastering video-text retrieval via image clip
Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen · 2021
Later among the works it cites.
Clip2tv: An empirical study on transformer-based methods for video-text retrieval
Zijian Gao, Jingyu Liu, Sheng Chen, Dedan Chang, Hao Zhang, and Jinwei Yuan · 2021
Later among the works it cites.
Ast: Audio spectrogram transformer
Yuan Gong, Yu-An Chung, and James Glass · 2021
Later among the works it cites.
Audio retrieval with natural language queries: A benchmark study
A Koepke, Andreea-Maria Oncescu, João F Henriques, Zeynep Akata, and Samuel Albanie · 2021
Later among the works it cites.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L. Berg, Mohit Bansal, and Jingjing Liu · 2021
Later among the works it cites.
Clip4clip: An empirical study of clip for end to end video clip retrieval
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li · 2021
Later among the works it cites.
Spoken moments: Learning joint audio-visual representations from video descriptions
Mathew Monfort, SouYoung Jin, Alexander Liu, David Harwath, Rogerio Feris, James Glass, and Aude Oliva · 2021
Later among the works it cites.
Attention bottlenecks for multimodal fusion
Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun · 2021
Later among the works it cites.
Audio retrieval with natural language queries
Andreea-Maria Oncescu, A Koepke, João F Henriques, Zeynep Akata, and Samuel Albanie · 2021
Later among the works it cites.
Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander G Hauptmann, Joao F. Henriques, and Andrea Vedaldi · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Look before you speak: Visually contextualized utterances
Paul Hongsuck Seo, Arsha Nagrani, and Cordelia Schmid · 2021
Later among the works it cites.
Decembert: Learning from noisy instructional videos via dense captions and entropy minimization
Zineng Tang, Jie Lei, and Mohit Bansal · 2021
Later among the works it cites.
T2vlad: Global-local sequence alignment for text-video retrieval, 2021
Xiaohan Wang, Linchao Zhu, and Yi Yang · 2021
Later among the works it cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Filip: Fine-grained interactive language-image pre-training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu · 2021
Later among the works it cites.
Everything at once–multi-modal fusion transformer for video retrieval
Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogerio Feris, David Harwath, James Glass, and Hilde Kuehne · 2022
Closest in time.