Fetching the paper…
Reading the bibliography…
Vision models trained on multimodal datasets can benefit from the wide availability of large image-caption datasets.
Big transfer (bit): General visual representation learning
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. 2019 · 1912
Earlier work this paper cites.
The generalization ofstudent’s’ problem when several different population variances are involved
Bernard L Welch. 1947 · 1947
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998 · 1998
Earlier work this paper cites.
WordNet: An electronic lexical database
George A Miller. 1998 · 1998
Earlier work this paper cites.
Placing search in context: the concept revisited
Lev Finkelstein, Evgeniy Gabrilovich, Yossi Matias, Ehud Rivlin, Zach Solan, Gadi Wolfman, and Eytan Ruppin. 2001 · 2001
Earlier work this paper cites.
Invariant visual representation by single neurons in the human brain
R Quian Quiroga, Leila Reddy, Gabriel Kreiman, Christof Koch, and Itzhak Fried. 2005 · 2005
Earlier work this paper cites.
Learning visual representations using images with captions
Ariadna Quattoni, Michael Collins, and Trevor Darrell. 2007 · 2007
Earlier work this paper cites.
Representational similarity analysis - connecting the branches of systems neuroscience
Nikolaus Kriegeskorte, Marieke Mur, and Peter Bandettini. 2008 · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al. 2009 · 2009
Earlier work this paper cites.
Software Framework for Topic Modelling with Large Corpora
Radim Řehůřek and Petr Sojka. 2010 · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Cited alongside, same era.
The Caltech-UCSD Birds-200-2011 Dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. 2011 · 2011
Cited alongside, same era.
Multimodal learning with deep boltzmann machines
Nitish Srivastava and Ruslan Salakhutdinov. 2012 · 2012
Cited alongside, same era.
Devise: A deep visual-semantic embedding model
Andrea Frome, Gregory S. Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomás Mikolov. 2013 · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Tomás Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013b · 2013
Cited alongside, same era.
Microsoft coco: Common objects in context
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. 2019 · 2019
Later among the works it cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019 · 2019
Later among the works it cites.
Found in translation: Learning robust joint representations by cyclic translations between modalities
Hai Pham, Paul Pu Liang, Thomas Manzini, Louis-Philippe Morency, and Barnabás Póczos. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Cited alongside, same era.
Concept cells through associative learning of high-level representations
Leila Reddy and Simon J Thorpe. 2014 · 2014
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Learning visual features from large weakly supervised data
Armand Joulin, Laurens Van Der Maaten, Allan Jabri, and Nicolas Vasilache. 2016 · 2016
Cited alongside, same era.
Foolbox: A python toolbox to benchmark the robustness of machine learning models
Jonas Rauber, Wieland Brendel, and Matthias Bethge. 2017 · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf. 2017 · 2017
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018 · 2018
Cited alongside, same era.
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman. 2020 · 2020
Later among the works it cites.
Do adversarially robust imagenet models transfer better?
Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor, and Aleksander Madry. 2020 · 2020
Later among the works it cites.
Learning visual representations with caption annotations
Mert Bulent Sariyildiz, Julien Perez, and Diane Larlus. 2020 · 2020
Later among the works it cites.
VirTex: Learning Visual Representations from Textual Annotations
Karan Desai and Justin Johnson. 2021 · 2021
Closest in time.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah. 2021 · 2021
Closest in time.
Transformer is all you need: Multimodal multitask learning with a unified transformer
Ronghang Hu and Amanpreet Singh. 2021 · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Closest in time.