Fetching the paper…
Reading the bibliography…
Traditional computer vision models are trained to predict a fixed set of predefined categories.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 1908
Earlier work this paper cites.
Image-to-word transformation based on dividing and vector quantizing images with words
Yasuhide Mori Hironobu, Hironobu Takahashi, and Ryuichi Oka · 1999
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning models for object recognition from natural language descriptions
Josiah Wang, Katja Markert, and Mark Everingham · 2009
Earlier work this paper cites.
Caltech-UCSD Birds 200
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona · 2010
Earlier work this paper cites.
Label-embedding for attribute-based classification
Zeynep Akata, Florent Perronnin, Zaid Harchaoui, and Cordelia Schmid · 2013
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S. Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
S. Maji, J. Kannala, E. Rahtu, M. Blaschko, and A. Vedaldi · 2013
Earlier work this paper cites.
Video event understanding using natural language descriptions
Vignesh Ramanathan, Percy Liang, and Li Fei-Fei · 2013
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li Fei-Fei · 2014
Earlier work this paper cites.
Zero-shot learning by convex combination of semantic embeddings
Mohammad Norouzi, Tomas Mikolov, Samy Bengio, Yoram Singer, Jonathon Shlens, Andrea Frome, Greg S. Corrado, and Jeffrey Dean · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
Richard Socher, Andrej Karpathy, Quoc V. Le, Christopher D. Manning, and Andrew Y. Ng · 2014
Earlier work this paper cites.
Evaluation of output embeddings for fine-grained image classification
Zeynep Akata, Scott Reed, Daniel Walter, Honglak Lee, and Bernt Schiele · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
An embarrassingly simple approach to zero-shot learning
Bernardino Romera-Paredes and Philip H. S. Torr · 2015
Earlier work this paper cites.
Optimal transport for domain adaptation
Nicolas Courty, Rémi Flamary, Devis Tuia, and Alain Rakotomamonjy · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Learning visual features from large weakly supervised data
Armand Joulin*, Laurens van der Maaten*, Allan Jabri, and Nicolas Vasilache · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Yfcc100m: The new data in multimedia research
Bart Thomee, David A. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Cited alongside, same era.
Learning visual n-grams from web data
Ang Li, Allan Jabri, Armand Joulin, and Laurens van der Maaten · 2017
Cited alongside, same era.
Label refinery: Improving ima- genet classification through label progression
wordnet: WordNet Interface , 2020
Ingo Feinerer and Kurt Hornik · 2020
Later among the works it cites.
Declutr: Deep contrastive learning for unsupervised textual representations
John M Giorgi, Osvald Nitski, Gary D. Bader, and Bo Wang · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, and Tom Duerig · 2020
Later among the works it cites.
Big transfer (bit): General visual representation learning
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hessam Bagherinezhad, Maxwell Horton, Mohammad Rastegari, and Ali Farhadi · 2018
Cited alongside, same era.
Bharath Bhushan Damodaran, Rémi Flamary, Vivien Seguy, and Nicolas Courty · 2018
Cited alongside, same era.
Improving gans using optimal transport
Tim Salimans, Han Zhang, Alec Radford, and Dimitris N. Metaxas · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Cited alongside, same era.
Zero-shot recognition via semantic embeddings and knowledge graphs
Xiaolong Wang, Yufei Ye, and Abhinav Gupta · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Cited alongside, same era.
Self-labelling via simultaneous clustering and representation learning
Yuki Markus Asano, Christian Rupprecht, and Andrea Vedaldi · 2019
Cited alongside, same era.
Later among the works it cites.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari · 2020
Later among the works it cites.
Hyperbolic visual embedding learning for zero-shot recognition
Shaoteng Liu, Jingjing Chen, Liangming Pan, Chong-Wah Ngo, Tat-Seng Chua, and Yu-Gang Jiang · 2020
Later among the works it cites.
S2ORC: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld · 2020
Later among the works it cites.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti · 2020
Later among the works it cites.
Contrastive learning with hard negative samples
Joshua Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka · 2020
Later among the works it cites.
Learning visual representations with caption annotations
Mert Bulent Sariyildiz, Julien Perez, and Diane Larlus · 2020
Later among the works it cites.
Learning from noisy labels with deep neural networks: A survey
Hwanjun Song, Minseok Kim, Dongmin Park, Yooju Shin, and Jae-Gil Lee · 2020
Later among the works it cites.
Fbnetv2: Differentiable neural architecture search for spatial and channel dimensions
Alvin Wan, Xiaoliang Dai, Peizhao Zhang, Zijian He, Yuandong Tian, Saining Xie, Bichen Wu, Matthew Yu, Tao Xu, Kan Chen, et al · 2020
Later among the works it cites.
Self-training with noisy student improves imagenet classification
Qizhe Xie, Eduard H. Hovy, Minh-Thang Luong, and Quoc V. Le · 2020
Later among the works it cites.
Contrastive learning of medical visual representations from paired images and texts
Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D. Manning, and Curtis P. Langlotz · 2020
Later among the works it cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Closest in time.
Unbiased teacher for semi-supervised object detection
Yen-Cheng Liu, Chih-Yao Ma, Zijian He, Chia-Wen Kuo, Kan Chen, Peizhao Zhang, Bichen Wu, Zsolt Kira, and Peter Vajda · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Closest in time.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork · 2021
Closest in time.