Fetching the paper…
Reading the bibliography…
Consider end-to-end training of a multi-modal vs.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Wsabie: Scaling up to large vocabulary image annotation
J. Weston, S. Bengio, and N. Usunier · 2011
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. A. Ranzato, and T. Mikolov · 2013
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
R. Socher, M. Ganjoo, C. D. Manning, and A. Y. Ng · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture
D. Eigen and R. Fergus · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Earlier work this paper cites.
Automatic description generation from images: A survey of models, datasets, and evaluation measures
R. Bernardi, R. Cakici, D. Elliott, A. Erdem, E. Erdem, N. Ikizler-Cinbis, F. Keller, A. Muscat, and B. Plank · 2016
Earlier work this paper cites.
Spatiotemporal residual networks for video action recognition
C. Feichtenhofer, A. Pinz, and R. P. Wildes · 2016
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. V. Gool · 2016
Earlier work this paper cites.
Actions ~ transformations
X. Wang, A. Farhadi, and A. Gupta · 2016
Earlier work this paper cites.
Yin and Yang: Balancing and answering binary visual questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Look, listen and learn
R. Arandjelović and A. Zisserman · 2017
Cited alongside, same era.
Gated multimodal units for information fusion
J. Arevalo, T. Solorio, M. M. y Gómez, and F. A. González · 2017
Cited alongside, same era.
Y. Bian, C. Gan, X. Liu, F. Li, X. Long, Y. Li, H. Qi, J. Zhou, S. Wen, and Y. Lin · 2017
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Efficient low-rank multimodal fusion with modality-specific factors
Z. Liu, Y. Shen, V. Lakshminarasimhan, P. Liang, A. Zadeh, and L.-P. Morency · 2018
Later among the works it cites.
Audio-visual scene analysis with self-supervised multisensory features
A. Owens and A. A. Efros · 2018
Later among the works it cites.
Hypothesis only baselines in natural language inference
A. Poliak, J. Naradowsky, A. Haldar, R. Rudinger, and B. Durme · 2018
Later among the works it cites.
Shifting the baseline: Single modality performance on visual navigation & qa
J. Thomason, D. Gordan, and Y. Bisk · 2018
Later among the works it cites.
A closer look at spatiotemporal convolutions for action recognition
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri · 2018
Later among the works it cites.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Cited alongside, same era.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Cited alongside, same era.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman · 2017
Cited alongside, same era.
Ubernet: Training a ‘universal’ convolutional neural network for low-, mid-, and high-level vision using diverse datasets and limited memory
I. Kokkinos · 2017
Cited alongside, same era.
Learning spatio-temporal representation with pseudo-3d residual networks
Z. Qiu, T. Yao, , and T. Mei · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Multimodal machine learning: A survey and taxonomy
T. Baltruvsaitis, C. Ahuja, and L.-P. Morency · 2018
Cited alongside, same era.
Later among the works it cites.
Y. Wang, J. Li, and F. Metze · 2018
Later among the works it cites.
Rethinking spatiotemporal feature learning for video understanding
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy · 2018
Later among the works it cites.
Multi-level attention model for weakly supervised audio classification
C. Yu, K. S. Barsim, Q. Kong, and B. Yang · 2018
Later among the works it cites.
The sound of pixels
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba · 2018
Later among the works it cites.
https://competitions.codalab.org/competitions/20115
Epic-kitchens action recognition · 2019
Closest in time.
Audio-visual scene-aware dialog
H. Alamri, V. Cartillier, A. Das, J. Wang, A. Cherian, I. Essa, D. B. amd Tim K. Marks, C. Hori, P. Anderson, S. Lee, and D. Parikh · 2019
Closest in time.
Slowfast networks for video recognition
C. Feichtenhofer, H. Fan, J. Malik, and K. He · 2019
Closest in time.
Large-scale weakly-supervised pre-training for video action recognition
D. Ghadiyaram, M. Feiszli, D. Tran, X. Yan, H. Wang, and D. K. Mahajan · 2019
Closest in time.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
E. Kazakos, A. Nagrani, A. Zisserman, and D. Damen · 2019
Closest in time.
Sgd on neural networks learns functions of increasing complexity
P. Nakkiran, G. Kaplun, D. Kalimeris, T. Yang, B. L. Edelman, F. Zhang, and B. Barak · 2019
Closest in time.
Video classification with channel-separated convolutional networks
D. Tran, H. Wang, L. Torresani, and M. Feiszli · 2019
Closest in time.
Baidu-uts submission to the epic-kitchens action recognition challenge 2019
X. Wang, Y. Wu, L. Zhu, and Y. Yang · 2019
Closest in time.
Imvotenet: Boosting 3d object detection in point clouds with image votes
C. R. Qi, X. Chen, O. Litany, and L. J. Guibas · 2020
Closest in time.