Fetching the paper…
Reading the bibliography…
Pre-trained representations are becoming crucial for many NLP and perception tasks.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Huang, Z., Zeng, Z., Liu, B., Fu, D., and Fu, J · 2004
Earlier work this paper cites.
M3p: Learning universal representations via multitask multilingual multimodal pre-training
Huang, H., Su, L., Qi, D., Duan, N., Cui, E., Bharti, T., Zhang, L., Wang, L., Gao, J., Liu, B., Fu, J., Zhang, D., Liu, X., and Zhou, M · 2006
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning the best pooling strategy for visual semantic embedding
Chen, J., Hu, H., Wu, H., Jiang, Y., and Wang, C · 2011
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. V · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Frome, A., Corrado, G. S., Shlens, J., Bengio, S., Dean, J., Ranzato, M. A., and Mikolov, T · 2013
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Food-101 – mining discriminative components with random forests
Bossard, L., Guillaumin, M., and Van Gool, L · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Karpathy, A., Joulin, A., and Li, F · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
Socher, R., Karpathy, A., Le, Q. V., Manning, C. D., and Ng, A. Y · 2014
Earlier work this paper cites.
Learning fine-grained image similarity with deep ranking
Wang, J., Song, Y., Leung, T., Rosenberg, C., Wang, J., Philbin, J., Chen, B., and Wu, Y · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollar, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
Simlex-999: Evaluating semantic models with (genuine) similarity estimation
Hill, F., Reichart, R., and Korhonen, A · 2015
Earlier work this paper cites.
Learning visual features from large weakly supervised data
Joulin, A., van der Maaten, L., Jabri, A., and Vasilache, N · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Multi30k: Multilingual english-german image descriptions
Elliott, D., Frank, S., Sima’an, K., and Specia, L · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., Bernstein, M., and Fei-Fei, L · 2016
Earlier work this paper cites.
Findings of the second shared task on multimodal machine translation and multilingual image description
Elliott, D., Frank, S., Barrault, L., Bougares, F., and Specia, L · 2017
Earlier work this paper cites.
Learning visual n-grams from web data
Li, A., Jabri, A., Joulin, A., and van der Maaten, L · 2017
Cited alongside, same era.
Dual attention networks for multimodal reasoning and matching
Nam, H., Ha, J.-W., and Kim, J · 2017
Cited alongside, same era.
Revisiting unreasonable effectiveness of data in deep learning era
Sun, C., Shrivastava, A., Sigh, S., and Gupta, A · 2017
Cited alongside, same era.
Findings of the third shared task on multimodal machine translation
Barrault, L., Bougares, F., Specia, L., Lala, C., Elliott, D., and Frank, S · 2018
Cited alongside, same era.
Vse++: Improving visual-semantic embeddings with hard negatives
Faghri, F., Fleet, D. J., Kiros, J. R., and Fidler, S · 2018
Cited alongside, same era.
Illustrative language understanding: Large-scale visual grounding with image search
Kiros, J., Chan, W., and Hinton, G · 2018
Graph-rise: Graph-regularized image semantic embedding
Juan, D.-C., Lu, C.-T., Li, Z., Peng, F., Timofeev, A., Chen, Y.-T., Gao, Y., Duerig, T., Tomkins, A., and Ravi, S · 2020
Later among the works it cites.
Big transfer (bit): General visual representation learning
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N · 2020
Later among the works it cites.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., Duerig, T., and Ferrari, V · 2020
Later among the works it cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., Choi, Y., and Gao, J · 2020
Later among the works it cites.
Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders
Messina, N., Amato, G., Esuli, A., Falchi, F., Gennaro, C., and Marchand-Maillet, S · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring the limits of weakly supervised pretraining
Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., and van der Maaten, L · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Visual semantic reasoning for image-text matching
Li, K., Zhang, Y., Li, K., Li, Y., and Fu, Y · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Self-supervised learning of pretext-invariant representations
Misra, I. and Maaten, L. v. d · 2020
Later among the works it cites.
A metric learning reality check
Musgrave, K., Belongie, S., and Lim, S.-N · 2020
Later among the works it cites.
Pham, H., Dai, Z., Xie, Q., Luong, M.-T., and Le, Q. V · 2020
Later among the works it cites.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Qi, D., Su, L., Song, J., Cui, E., Bharti, T., and Sacheti, A · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Later among the works it cites.
Learning visual representations with caption annotations
Sariyildiz, M. B., Perez, J., and Larlus, D · 2020
Later among the works it cites.
Contrastive multiview coding
Tian, Y., Krishnan, D., and Isola, P · 2020
Later among the works it cites.
Self-training with noisy student improves imagenet classification
Xie, Q., Luong, M.-T., Hovy, E., and Le, Q. V · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2020
Later among the works it cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
Yu, F., Tang, J., Yin, W., Sun, Y., Tian, H., Wu, H., and Wang, H · 2020
Later among the works it cites.
Contrastive learning of medical visual representations from paired images and text
Zhang, Y., Jiang, H., Miura, Y., Manning, C. D., and Langlotz, C. P · 2020
Later among the works it cites.
Rethinking pre-training and self-training
Zoph, B., Ghiasi, G., Lin, T.-Y., Cui, Y., Liu, H., Cubuk, E. D., and Le, Q. V · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Closest in time.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Closest in time.
Natural adversarial examples
Hendrycks, D., Zhao, K., Basart, S., Steinhardt, J., and Song, D · 2021
Closest in time.
Prototypical contrastive learning of unsupervised representations
Li, J., Zhou, P., Xiong, C., and Hoi, S · 2021
Closest in time.
Crisscrossed captions: Extended intramodal and intermodal semantic similarity judgments for ms-coco
Parekh, Z., Baldridge, J., Cer, D., Waters, A., and Yang, Y · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarawl, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Closest in time.
UC2: Universal cross-lingual cross-modal vision-and-language pre-training
Zhou, M., Zhou, L., Wang, S., Cheng, Y., Li, L., Yu, Z., and Liu, J · 2021
Closest in time.