Fetching the paper…
Reading the bibliography…
Large Vision & Language models pretrained on web-scale data provide representations that are invaluable for numerous V&L problems.
Carey, S., Bartlett, E.: Acquiring a single new word. (1978)
1978
Earlier work this paper cites.
Markman, E.M.: Constraints children place on word meanings. Cognitive science 14
1990
Earlier work this paper cites.
Markman, E.M., Wasow, J.L., Hansen, M.B.: Use of the mutual exclusivity assumption by young word learners. Cognitive psychology 47
2003
Earlier work this paper cites.
Dekel, O., Keshet, J., Singer, Y.: Large margin hierarchical classification. In: Proceedings of the twenty-first international conference on Machine learning. p. 27 (2004)
2004
Earlier work this paper cites.
Lampert, C., Nickisch, H., Harmeling, S.: Learning to detect unseen object classes by between-class attribute transfer. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2009)
2009
Earlier work this paper cites.
2014
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: European conference on computer vision. pp. 740–755. Springer (2014)
2014
Earlier work this paper cites.
Denton, E., Weston, J., Paluri, M., Bourdev, L., Fergus, R.: User conditional hashtag prediction for images. In: Proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining. pp. 1731–1740 (2015)
2015
Earlier work this paper cites.
Akata, Z., Perronnin, F., Harchaoui, Z., Schmid, C.: Label-embedding for image classification. IEEE Transactions on Pattern Analysis and Machine Intelligence 38
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Hendricks, L.A., Venugopalan, S., Rohrbach, M., Mooney, R., Saenko, K., Darrell, T.: Deep compositional captioning: Describing novel object categories without paired training data. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 1–10 (2016)
2016
Earlier work this paper cites.
Vinyals, O., Blundell, C., Lillicrap, T., Wierstra, D., et al.: Matching networks for one shot learning. Advances in neural information processing systems 29
2016
Earlier work this paper cites.
Chunseong Park, C., Kim, B., Kim, G.: Attend to you: Personalized image captioning with context sequence memory networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 895–903 (2017)
2017
Earlier work this paper cites.
Feng, F., Liu, R., Wang, X., Li, X., Bi, S.: Personalized image annotation using deep architecture. IEEE Access 5
2017
Earlier work this paper cites.
Finn, C., Abbeel, P., Levine, S.: Model-agnostic meta-learning for fast adaptation of deep networks. In: International conference on machine learning. pp. 1126–1135. PMLR (2017)
2017
Earlier work this paper cites.
Snell, J., Swersky, K., Zemel, R.: Prototypical networks for few-shot learning. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
Venugopalan, S., Anne Hendricks, L., Rohrbach, M., Mooney, R., Darrell, T., Saenko, K.: Captioning images with diverse objects. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 5753–5761 (2017)
2017
Earlier work this paper cites.
Xian, Y., Schiele, B., Akata, Z.: Zero-shot learning - the good, the bad and the ugly. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2017)
2017
Earlier work this paper cites.
Zaheer, M., Kottur, S., Ravanbakhsh, S., Poczos, B., Salakhutdinov, R.R., Smola, A.J.: Deep sets. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
Atzmon, Y., Chechik, G.: Probabilistic and-or attribute grouping for zero-shot learning. In: Proceedings of the Thirty-Forth Conference on Uncertainty in Artificial Intelligence (2018)
2018
Earlier work this paper cites.
Lu, J., Yang, J., Batra, D., Parikh, D.: Neural baby talk. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 7219–7228 (2018)
2018
Earlier work this paper cites.
Ren, M., Ravi, S., Triantafillou, E., Snell, J., Swersky, K., Tenenbaum, J.B., Larochelle, H., Zemel, R.S.: Meta-learning for semi-supervised few-shot classification. In: International Conference on Learning Representations (2018), https://openreview.net/forum?id=HJcSzz-CZ
2018
Earlier work this paper cites.
Sung, F., Yang, Y., Zhang, L., Xiang, T., Torr, P.H., Hospedales, T.M.: Learning to compare: Relation network for few-shot learning. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 1199–1208 (2018)
2018
Earlier work this paper cites.
Wu, Y., Zhu, L., Jiang, L., Yang, Y.: Decoupled novel object captioner. In: Proceedings of the 26th ACM international conference on Multimedia. pp. 1029–1037 (2018)
2018
Earlier work this paper cites.
Xu, N., Yang, L., Fan, Y., Yang, J., Yue, D., Liang, Y., Price, B., Cohen, S., Huang, T.: Youtube-vos: Sequence-to-sequence video object segmentation. In: Proceedings of the European conference on computer vision (ECCV). pp. 585–601 (2018)
2018
Earlier work this paper cites.
Abdal, R., Qin, Y., Wonka, P.: Image2stylegan: How to embed images into the stylegan latent space? In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 4432–4441 (2019)
2019
Earlier work this paper cites.
Atzmon, Y., Chechik, G.: Adaptive confidence smoothing for generalized zero-shot learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 11671–11680 (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Ge, Y., Zhang, R., Wu, L., Wang, X., Tang, X., Luo, P.: A versatile benchmark for detection, pose estimation, segmentation and re-identification of clothing images. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Earlier work this paper cites.
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., Gelly, S.: Parameter-efficient transfer learning for nlp. In: International Conference on Machine Learning. pp. 2790–2799. PMLR (2019)
2019
Earlier work this paper cites.
Hsieh, Y.G., Niu, G., Sugiyama, M.: Classification from positive, unlabeled and biased negative data. In: International Conference on Machine Learning. pp. 2820–2829. PMLR (2019)
2019
Cited alongside, same era.
Shuster, K., Humeau, S., Hu, H., Bordes, A., Weston, J.: Engaging image captioning via personality. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 12516–12526 (2019)
2019
Cited alongside, same era.
Zheng, Y., Li, Y., Wang, S.: Intention oriented image captions with guiding objects. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8395–8404 (2019)
2019
Cited alongside, same era.
Chen, T., Kornblith, S., Norouzi, M., Hinton, G.: A simple framework for contrastive learning of visual representations. In: International conference on machine learning. pp. 1597–1607. PMLR (2020)
2020
Cited alongside, same era.
2021
Later among the works it cites.
Liu, N., Li, S., Du, Y., Tenenbaum, J., Torralba, A.: Learning to compose visual relations. Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, Y., Gong, S., Bazzani, L.: Image search with text feedback by visiolinguistic attention learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 3001–3011 (2020)
2020
Cited alongside, same era.
Del Chiaro, R., Twardowski, B., Bagdanov, A., Van de Weijer, J.: Ratt: Recurrent attention to transient tasks for continual image captioning. Advances in Neural Information Processing Systems 33
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Jia, X., Zhao, H., Lin, Z., Kale, A., Kumar, V.: Personalized image retrieval with sparse graph representation learning. In: Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. pp. 2735–2743 (2020)
2020
Cited alongside, same era.
Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., et al.: The open images dataset v4. International Journal of Computer Vision 128
2020
Cited alongside, same era.
Lake, B.M., Piantadosi, S.T.: People infer recursive visual concepts from just a few examples. Computational Brain & Behavior 3
2020
Cited alongside, same era.
Long, C., Yang, X., Xu, C.: Cross-domain personalized image captioning. Multimedia Tools and Applications 79
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Patashnik, O., Wu, Z., Shechtman, E., Cohen-Or, D., Lischinski, D.: Styleclip: Text-driven manipulation of stylegan imagery. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 2085–2094 (2021)
2021
Later among the works it cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning. pp. 8748–8763. PMLR (2021)
2021
Later among the works it cites.
Richardson, E., Alaluf, Y., Patashnik, O., Nitzan, Y., Azar, Y., Shapiro, S., Cohen-Or, D.: Encoding in style: a stylegan encoder for image-to-image translation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 2287–2296 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Shridhar, M., Manuelli, L., Fox, D.: Cliport: What and where pathways for robotic manipulation. In: Proceedings of the 5th Conference on Robot Learning (CoRL) (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Tsimpoukelli, M., Menick, J., Cabi, S., Eslami, S., Vinyals, O., Hill, F.: Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
2021
Later among the works it cites.
Wu, G., Gong, S., Li, P.: Striking a balance between stability and plasticity for class-incremental learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 1124–1133 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Zabari, N., Hoshen, Y.: Semantic segmentation in-the-wild without seeing any segmentation examples (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Zhang, Y., Zhang, C.B., Jiang, P.T., Cheng, M.M., Mao, F.: Personalized image semantic segmentation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 10549–10559 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Alaluf, Y., Tov, O., Mokady, R., Gal, R., Bermano, A.: Hyperstyle: Stylegan inversion with hypernetworks for real image editing. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 18511–18521 (2022)
2022
Closest in time.
Dinh, T.M., Tran, A.T., Nguyen, R., Hua, B.S.: Hyperinverter: Improving stylegan inversion via hypernetwork. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 11389–11398 (2022)
2022
Closest in time.
2022
Closest in time.
Li, B., Weinberger, K.Q., Belongie, S., Koltun, V., Ranftl, R.: Language-driven semantic segmentation. In: International Conference on Learning Representations (2022), https://openreview.net/forum?id=RriDjddCLN
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
Wang, L., Meng, X., Xiang, Y., Fox, D.: Hierarchical policies for cluttered-scene grasping with latent plans. IEEE Robotics and Automation Letters (2022)
2022
Closest in time.