Fetching the paper…
Reading the bibliography…
The fundamental goal of artificial intelligence (AI) is to mimic the core cognitive activities of human.
Invariant visual representation by single neurons in the human brain
Quiroga, R. Q., Reddy, L., Kreiman, G., Koch, C. & Fried, I · 2005
Earlier work this paper cites.
A comparison and semi-quantitative analysis of words and character-bigrams as features in chinese text categorization
Li, J., Sun, M. & Zhang, X · 2006
Earlier work this paper cites.
Explicit encoding of multimodal percepts by single neurons in the human brain
Quian Quiroga, R., Kraskov, A., Koch, C. & Fried, I · 2009
Earlier work this paper cites.
Bag-of-visual-words and spatial extensions for land-use classification
Yang, Y. & Newsam, S. D · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. & Hinton, G. E · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G. & Berg, T. L · 2011
Earlier work this paper cites.
Artificial general intelligence: Concept, state of the art, and future prospects
Goertzel, B · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y. et al · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I. J. et al · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M. & Hockenmaier, J · 2014
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y. & Hinton, G · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Russakovsky, O. et al · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R. B. & Sun, J · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. & Ba, J · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S. & Sun, J · 2016
Earlier work this paper cites.
Visual7W: Grounded question answering in images
Zhu, Y., Groth, O., Bernstein, M. S. & Fei-Fei, L · 2016
Earlier work this paper cites.
A simple neural network module for relational reasoning
Santoro, A. et al · 2017
Earlier work this paper cites.
Feature visualization
Olah, C., Mordvintsev, A. & Schubert, L · 2017
Earlier work this paper cites.
AID: A benchmark data set for performance evaluation of aerial scene classification
Xia, G.-S. et al · 2017
Earlier work this paper cites.
Zero-shot scene classification for high spatial resolution remote sensing images
Li, A., Lu, Z., Wang, L., Xiang, T. & Wen, J · 2017
Earlier work this paper cites.
AI challenger: A large-scale dataset for going deeper in image understanding
Wu, J. et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A. et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
van den Oord, A., Vinyals, O. & Kavukcuoglu, K · 2017
Cited alongside, same era.
Unsupervised feature learning via non-parametric instance discrimination
Wu, Z., Xiong, Y., Yu, S. X. & Lin, D · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y. & Vinyals, O · 2018
Cited alongside, same era.
UMAP: Uniform manifold approximation and projection
McInnes, L., Healy, J., Saul, N. & Großberger, L · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T. & Sutskever, I · 2018
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S. & Girshick, R · 2020
Later among the works it cites.
Bootstrap your own latent - a new approach to self-supervised learning
Grill, J.-B. et al · 2020
Later among the works it cites.
Revisiting pre-trained models for chinese natural language processing
Cui, Y. et al · 2020
Later among the works it cites.
“Jieba” Chinese text segmentation
Sun, J · 2020
Later among the works it cites.
Big Transfer (BiT): General visual representation learning
Kolesnikov, A. et al · 2020
Later among the works it cites.
Multi-modality cross attention network for image and sentence matching
Wei, X., Zhang, T., Li, Y., Zhang, Y. & Wu, F · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anderson, P. et al · 2018
Cited alongside, same era.
RoBERTa: A robustly optimized bert pretraining approach
Liu, Y. et al · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A. et al · 2019
Cited alongside, same era.
VisualBERT: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C. & Chang, K · 2019
Cited alongside, same era.
Learning fragment self-attention embeddings for image-text matching
Wu, Y., Wang, S., Song, G. & Huang, Q · 2019
Cited alongside, same era.
Learning deep representations by mutual information estimation and maximization
Hjelm, R. D. et al · 2019
Cited alongside, same era.
Local aggregation for unsupervised learning of visual embeddings
Zhuang, C., Zhai, A. L. & Yamins, D · 2019
Cited alongside, same era.
Bommasani, R. et al · 2021
Closest in time.
10 breakthrough technologies 2021
Editors of MIT Technology Review · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Radford, A. et al · 2021
Closest in time.
M6: A chinese multimodal pretrainer
Lin, J. et al · 2021
Closest in time.
Similarity reasoning and filtration for image-text matching
Diao, H., Zhang, Y., Ma, L. & Lu, H · 2021
Closest in time.
ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learning
Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S. & He, Y · 2021
Closest in time.
Scaling vision with sparse mixture of experts
Riquelme, C. et al · 2021
Closest in time.
Exploring simple siamese representation learning
Chen, X. & He, K · 2021
Closest in time.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C. et al · 2021
Closest in time.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R. & Ommer, B · 2021
Closest in time.
Zero and few shot learning with semantic feature synthesis and competitive learning
Guan, J. et al · 2021
Closest in time.
Toutiao text classification dataset
Contributors of Toutiao News · 2021
Closest in time.
Counterfactual VQA: A cause-effect look at language bias
Niu, Y. et al · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A. et al · 2021
Closest in time.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S. & Soricut, R · 2021
Closest in time.