Fetching the paper…
Reading the bibliography…
Vision-language models are pre-trained by aligning image-text pairs in a common space to deal with open-set visual concepts.
C. Granger, “Investigating causal relations by econometric models and cross-spectral methods,” Econometrica , vol. 37, pp. 424–38, 02 1969
1969
Earlier work this paper cites.
Granger and C. W., “Testing for causality: a personal viewpoint,” in Journal of Economic Dynamics and Control, 2:329–352, 1980 , 1980
1980
Earlier work this paper cites.
T. K. Nakamura, “Statistical mechanics of a collisionless system based on the maximum entropy principle,” The Astrophysical Journal , vol. 531, no. 2, p. 739, 2000
2000
Earlier work this paper cites.
F. Li, R. Fergus, and P. Perona, “Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,” in Conference on Computer Vision & Pattern Recognition Workshop , 2004
2004
Earlier work this paper cites.
M. Yi, D. Harm, H. Wei, and W. John, “Segmentation of multivariate mixed data via lossy data coding and compression,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 29, no. 9, pp. 1546–1562, 2007. [Online]. Available: https://doi.org/10.1109/TPAMI.2007.1085
2007
Earlier work this paper cites.
Nilsback, ME, and Zisserman, “Automated flower classification over a large number of classes,” ICVGIP , pp. 722–729, 2008
2008
Earlier work this paper cites.
J. Tang, J. Zhang, L. Yao, J. Li, L. Zhang, and Z. Su, “Arnetminer: extraction and mining of academic social networks,” in KDD . ACM, 2008, pp. 990–998
2008
Earlier work this paper cites.
J. Pearl, “Causal inference in statistics: An overview,” Statistics surveys , pp. 96–146, 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR . IEEE Computer Society, 2009, pp. 248–255
2009
Earlier work this paper cites.
A. Carlson, J. Betteridge, B. Kisiel, B. Settles, E. R. H. Jr., and T. M. Mitchell, “Toward an architecture for never-ending language learning,” in AAAI . AAAI Press, 2010
2010
Earlier work this paper cites.
D. A. Ferrucci, E. W. Brown, J. Chu-Carroll, J. Fan, D. Gondek, A. Kalyanpur, A. Lally, J. W. Murdock, E. Nyberg, J. M. Prager, N. Schlaefer, and C. A. Welty, “Building watson: An overview of the deepqa project,” AI Mag. , vol. 31, no. 3, pp. 59–79, 2010
2010
Earlier work this paper cites.
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba, “SUN database: Large-scale scene recognition from abbey to zoo,” in CVPR . IEEE Computer Society, 2010, pp. 3485–3492
2010
Earlier work this paper cites.
M. Nickel, V. Tresp, and H. Kriegel, “A three-way model for collective learning on multi-relational data,” in ICML . Omnipress, 2011, pp. 809–816
2011
Earlier work this paper cites.
L. Harland, “Open PHACTS: A semantic knowledge infrastructure for public and commercial drug discovery research,” in EKAW , ser. Lecture Notes in Computer Science, vol. 7603. Springer, 2012, pp. 1–7
2012
Earlier work this paper cites.
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. V. Jawahar, “Cats and dogs,” in CVPR . IEEE Computer Society, 2012, pp. 3498–3505
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
X. Glorot, A. Bordes, J. Weston, and Y. Bengio, “A semantic matching energy function for learning with multi-relational data,” in ICLR (Workshop Poster) , 2013
2013
Earlier work this paper cites.
R. Socher, D. Chen, C. D. Manning, and A. Y. Ng, “Reasoning with neural tensor networks for knowledge base completion,” in NIPS , 2013, pp. 926–934
2013
Earlier work this paper cites.
A. Bordes, N. Usunier, A. García-Durán, J. Weston, and O. Yakhnenko, “Translating embeddings for modeling multi-relational data,” in NIPS , 2013, pp. 2787–2795
2013
Earlier work this paper cites.
J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3d object representations for fine-grained categorization,” in ICCV Workshops . IEEE Computer Society, 2013, pp. 554–561
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
D. Vrandecic and M. Krötzsch, “Wikidata: a free collaborative knowledgebase,” Commun. ACM , vol. 57, no. 10, pp. 78–85, 2014
2014
Earlier work this paper cites.
L. Bossard, M. Guillaumin, and L. V. Gool, “Food-101 - mining discriminative components with random forests,” in ECCV (6) , ser. Lecture Notes in Computer Science. Springer, 2014
2014
Earlier work this paper cites.
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi, “Describing textures in the wild,” in CVPR . IEEE Computer Society, 2014, pp. 3606–3613
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: visual question answering,” in ICCV . IEEE Computer Society, 2015, pp. 2425–2433
2015
Cited alongside, same era.
Y. Lin, Z. Liu, M. Sun, Y. Liu, and X. Zhu, “Learning entity and relation embeddings for knowledge graph completion,” in AAAI . AAAI Press, 2015, pp. 2181–2187
2015
Cited alongside, same era.
A. Goldstein, A. Kapelner, J. Bleich, and E. Pitkin, “Peeking inside the black box: Visualizing statistical learning with plots of individual conditional expectation.” in Journal of Computational and Graphical Statistics, 24(1):44–65, 2015. , 2015
2015
Cited alongside, same era.
M. Glymour, J. Pearl, and N. P. Jewell, Causal inference in statistics: A primer . John Wiley & Sons, 2016
2016
Cited alongside, same era.
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo, “Image captioning with semantic attention,” in CVPR . IEEE Computer Society, 2016, pp. 4651–4659
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar, “Do imagenet classifiers generalize to imagenet?” in ICML , ser. Proceedings of Machine Learning Research, vol. 97. PMLR, 2019, pp. 5389–5400
2019
Later among the works it cites.
H. Wang, S. Ge, Z. C. Lipton, and E. P. Xing, “Learning robust global representations by penalizing local predictive power,” in NeurIPS , 2019, pp. 10 506–10 518
2019
Later among the works it cites.
T. Shin, Y. Razeghi, R. L. L. IV, E. Wallace, and S. Singh, “Autoprompt: Eliciting knowledge from language models with automatically generated prompts,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 , 2020
2020
Later among the works it cites.
Z. Jiang, F. F. Xu, J. Araki, and G. Neubig, “How can we know what language models know,” Trans. Assoc. Comput. Linguistics , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016 . IEEE Computer Society, 2016, pp. 770–778
2016
Cited alongside, same era.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2016
2016
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA , 2017
2017
Cited alongside, same era.
B. Xu, Y. Xu, J. Liang, C. Xie, B. Liang, W. Cui, and Y. Xiao, “Cn-dbpedia: A never-ending chinese knowledge extraction system,” in IEA/AIE (2) , ser. Lecture Notes in Computer Science, vol. 10351. Springer, 2017, pp. 428–438
2017
Cited alongside, same era.
R. Speer, J. Chin, and C. Havasi, “Conceptnet 5.5: An open multilingual graph of general knowledge,” in AAAI . AAAI Press, 2017, pp. 4444–4451
2017
Cited alongside, same era.
Q. Wang, Z. Mao, B. Wang, and L. Guo, “Knowledge graph embedding: A survey of approaches and applications,” IEEE Trans. Knowl. Data Eng. , vol. 29, no. 12, pp. 2724–2743, 2017
2017
Cited alongside, same era.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in CVPR . Computer Vision Foundation / IEEE Computer Society, 2018, pp. 6077–6086
2018
Cited alongside, same era.
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, Y. Choi, and J. Gao, “Oscar: Object-semantics aligned pre-training for vision-language tasks,” in ECCV (30) , ser. Lecture Notes in Computer Science, vol. 12375. Springer, 2020, pp. 121–137
2020
Later among the works it cites.
2020
Later among the works it cites.
P. Qin, X. Wang, W. Chen, C. Zhang, W. Xu, and W. Y. Wang, “Generative adversarial zero-shot relational learning for knowledge graphs,” in AAAI . AAAI Press, 2020, pp. 8673–8680
2020
Later among the works it cites.
T. Chen, S. Kornblith, M. Norouzi, and G. E. Hinton, “A simple framework for contrastive learning of visual representations,” in Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event , ser. Proceedings of Machine Learning Research. PMLR, 2020
2020
Later among the works it cites.
Y. Tian, Y. Wang, D. Krishnan, J. B. Tenenbaum, and P. Isola, “Rethinking few-shot image classification: A good embedding is all you need?” in Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XIV , ser. Lecture Notes in Computer Science, 2020
2020
Later among the works it cites.
T. Schick and H. Schütze, “Exploiting cloze-questions for few-shot text classification and natural language inference,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL 2021, Online, April 19 - 23, 2021 , 2021
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event , 2021
2021
Later among the works it cites.
Y. Rao, W. Zhao, G. Chen, Y. Tang, Z. Zhu, G. Huang, J. Zhou, and J. Lu, “Denseclip: Language-guided dense prediction with context-aware prompting,” CoRR , 2021
2021
Later among the works it cites.
P. Gao, S. Geng, R. Zhang, T. Ma, R. Fang, Y. Zhang, H. Li, and Y. Qiao, “Clip-adapter: Better vision-language models with feature adapters,” CoRR , 2021
2021
Later among the works it cites.
C. Jia, Y. Yang, Y. Xia, Y. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event , 2021
2021
Later among the works it cites.
W. Lin, H. Lan, and B. Li, “Generative causal explanations for graph neural networks,” in Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event , ser. Proceedings of Machine Learning Research. PMLR, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Geng, J. Chen, Z. Chen, J. Z. Pan, Z. Ye, Z. Yuan, Y. Jia, and H. Chen, “Ontozsl: Ontology-enhanced zero-shot learning,” in WWW . ACM / IW3C2, 2021, pp. 3325–3336
2021
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 , 2021
2021
Later among the works it cites.
D. Hendrycks, K. Zhao, S. Basart, J. Steinhardt, and D. Song, “Natural adversarial examples,” in CVPR . Computer Vision Foundation / IEEE, 2021, pp. 15 262–15 271
2021
Later among the works it cites.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision-language models,” Int. J. Comput. Vis. , vol. 130, no. 9, pp. 2337–2348, 2022. [Online]. Available: https://doi.org/10.1007/s11263-022-01653-1
2022
Closest in time.
W. Jin, Y. Cheng, Y. Shen, W. Chen, and X. Ren, “A good prompt is worth millions of parameters: Low-resource prompt-based learning for vision-language models,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022 , 2022
2022
Closest in time.
2022
Closest in time.
X. Liu, Z. Wang, Y.-L. Li, and S. Wang, “Self-supervised learning via maximum entropy coding,” in Advances in Neural Information Processing Systems , A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, Eds., 2022. [Online]. Available: https://openreview.net/forum?id=nJt27NQffr
2022
Closest in time.
J. Yu, H. Yin, X. Xia, T. Chen, L. Cui, and Q. V. H. Nguyen, “Are graph augmentations necessary?: Simple graph contrastive learning for recommendation,” in SIGIR ’22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, July 11 - 15, 2022 , E. Amigó, P. Castells, J. Gonzalo, B. Carterette, J. S. Culpepper, and G. Kazai, Eds. ACM, 2022, pp. 1294–1303. [Online]. Available: https://doi.org/10.1145/3477495.3531937
2022
Closest in time.