Fetching the paper…
Reading the bibliography…
Many real-world video-text tasks involve different levels of granularity, such as frames and words, clip and sentences or videos and paragraphs, each with distinct semantics.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Neighbourhood components analysis
Jacob Goldberger, Geoffrey E Hinton, Sam T. Roweis, and Russ R Salakhutdinov · 2005
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
R. Hadsell, S. Chopra, and Y. LeCun · 2006
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Learnable pooling regions for image classification
Mateusz Malinowski and Mario Fritz · 2013
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc' Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li F Fei-Fei · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Visual word2vec (vis-w2v): Learning visually grounded word embeddings using abstract scenes
Satwik Kottur, Ramakrishna Vedantam, José M. F. Moura, and Devi Parikh · 2015
Earlier work this paper cites.
A hierarchical neural autoencoder for paragraphs and documents
Jiwei Li, Thang Luong, and Dan Jurafsky · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using fisher vectors
B. Klein, G. Lev, G. Sadeh, and L. Wolf · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Sequence to sequence-video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Learning deep structure-preserving image-text embeddings
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Jointly modeling embedding and translation to bridge video and language
Yingwei Pan, Tao Mei, Ting Yao, Houqiang Li, and Yong Rui · 2015
Earlier work this paper cites.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
Ran Xu, Caiming Xiong, Wei Chen, and Jason J. Corso · 2015
Earlier work this paper cites.
Weakly-supervised alignment of video with text
Piotr Bojanowski, Rémi Lajugie, Edouard Grave, Francis Bach, Ivan Laptev, Jean Ponce, and Cordelia Schmid · 2015
Earlier work this paper cites.
Book2movie: Aligning video scenes with book chapters
M. Tapaswi, M. Bäuml, and R. Stiefelhagen · 2015
Earlier work this paper cites.
Grounding of textual phrases in images by reconstruction
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele · 2016
Earlier work this paper cites.
Hierarchical recurrent neural encoder for video representation with application to captioning
P. Pan, Z. Xu, Y. Yang, F. Wu, and Y. Zhuang · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Dual learning for machine translation
Di He, Yingce Xia, Tao Qin, Liwei Wang, Nenghai Yu, Tie-Yan Liu, and Wei-Ying Ma · 2016
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2016
Earlier work this paper cites.
Enhancing video summarization via vision-language embedding
B. A. Plummer, M. Brown, and S. Lazebnik · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
No fuss distance metric learning using proxies
Yair Movshovitz-Attias, Alexander Toshev, Thomas K. Leung, Sergey Ioffe, and Saurabh Singh · 2017
Cited alongside, same era.
Near-duplicate video retrieval with deep metric learning
G. Kordopatis-Zilo, S. Papadopoulos, I. Patras, and Y. Kompatsiaris · 2017
Cited alongside, same era.
Learnable pooling with context gating for video classification
Antoine Miech, Ivan Laptev, and Josef Sivic · 2017
Cited alongside, same era.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2019
Later among the works it cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
Phrase localization and visual relationship detection with comprehensive image-language cues
Bryan A Plummer, Arun Mallya, Christopher M Cervantes, Julia Hockenmaier, and Svetlana Lazebnik · 2017
Cited alongside, same era.
See, hear, and read: Deep aligned representations
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2017
Cited alongside, same era.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Cited alongside, same era.
TALL: temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Cited alongside, same era.
Unpaired image-to-image translation using cycle-consistent adversarial networks, 2017
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros · 2017
Cited alongside, same era.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Cited alongside, same era.
Later among the works it cites.
Use what you have: Video retrieval using representations from collaborative experts
Y. Liu, S. Albanie, A. Nagrani, and A. Zisserman · 2019
Later among the works it cites.
Contrastive bidirectional transformer for temporal representation learning
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
Mule: Multimodal universal language embedding, 2019
Donghyun Kim, Kuniaki Saito, Kate Saenko, Stan Sclaroff, and Bryan A. Plummer · 2019
Later among the works it cites.
Language models are unsupervised multitask learners, 2019
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Temporal cycle-consistency learning
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, and Andrew Zisserman · 2019
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context, 2019
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Joint visual-textual embedding for multimodal style search
Gil Sadeh, Lior Fritz, Gabi Shalev, and Eduard Oks · 2019
Later among the works it cites.
Uniter: Universal image-text representation learning, 2019
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2019
Later among the works it cites.
Unified vision-language pre-training for image captioning and vqa
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao · 2019
Later among the works it cites.
Dual encoding for zero-example video retrieval
Jianfeng Dong, Xirong Li, Chaoxi Xu, Shouling Ji, Yuan He, Gang Yang, and Xun Wang · 2019
Later among the works it cites.
Wslln:weakly supervised natural language localization networks
Mingfei Gao, Larry Davis, Richard Socher, and Caiming Xiong · 2019
Later among the works it cites.
Learning correspondence from the cycle-consistency of time
Xiaolong Wang, Allan Jabri, and Alexei A. Efros · 2019
Later among the works it cites.
Cycle-consistency for robust visual question answering
Meet Shah, Xinlei Chen, Marcus Rohrbach, and Devi Parikh · 2019
Later among the works it cites.
Crdoco: Pixel-level domain transfer with cross-domain consistency
Y. Chen, Y. Lin, M. Yang, and J. Huang · 2019
Later among the works it cites.
Look closer to ground better: Weakly-supervised temporal grounding of sentence in video
Zhenfang Chen, Lin Ma, Wenhan Luo, Peng Tang, and Kwan-Yee K. Wong · 2020
Closest in time.
Fine-grained video-text retrieval with hierarchical graph reasoning
Shizhe Chen, Yida Zhao, Qin Jin, and Qi Wu · 2020
Closest in time.
Actbert: Learning global-local video-text representations
Yi Yang Linchao Zhu · 2020
Closest in time.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Closest in time.
Mart: Memory-augmented recurrent transformer for coherent video paragraph captioning
Jie Lei, Liwei Wang, Yelong Shen, Dong Yu, Tamara L Berg, and Mohit Bansal · 2020
Closest in time.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data, 2020
Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti · 2020
Closest in time.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers, 2020
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu · 2020
Closest in time.
lambert: Language and action learning using multimodal bert
Kazuki Miyazawa, Tatsuya Aoki, Takato Horii, and Takayuki Nagai · 2020
Closest in time.
Local-global video-text interactions for temporal grounding
Jonghwan Mun, Minsu Cho, , and Bohyung Han · 2020
Closest in time.
Multi-modal transformer for video retrieval, 2020
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Closest in time.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2020
Closest in time.