Fetching the paper…
Reading the bibliography…
Many high-level skills that are required for computer vision tasks, such as parsing questions, comparing and contrasting semantics, and writing descriptions, are also required in other domains such as natural language processing.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Integrating structured biological data by kernel maximum mean discrepancy
Karsten M Borgwardt, Arthur Gretton, Malte J Rasch, Hans-Peter Kriegel, Bernhard Schölkopf, and Alex J Smola · 2006
Earlier work this paper cites.
Modeling semantic containment and exclusion in natural language inference
Bill MacCartney and Christopher D. Manning · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Frustratingly easy domain adaptation
Hal Daumé III · 2009
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Deep domain confusion: Maximizing for domain invariance
Eric Tzeng, Judy Hoffman, Ning Zhang, Kate Saenko, and Trevor Darrell · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky · 2016
Earlier work this paper cites.
Senticap: Generating image descriptions with sentiments
A. Mathews, Lexing Xie, and Xuming He · 2016
Earlier work this paper cites.
Show, adapt and tell: Adversarial training of cross-domain image captioner
Tseng-Hung Chen, Yuan-Hong Liao, Ching-Yao Chuang, Wan Ting Hsu, Jianlong Fu, and Min Sun · 2017
Earlier work this paper cites.
Stylenet: Generating attractive visual captions with styles
Chuang Gan, Zhe Gan, Xiaodong He, Jianfeng Gao, and Li Deng · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Lukasz Kaiser, Aidan N. Gomez, Noam M. Shazeer, Ashish Vaswani, Niki Parmar, Llion Jones, and Jakob Uszkoreit · 2017
Earlier work this paper cites.
Adversarial discriminative domain adaptation
Eric Tzeng, Judy Hoffman, Kate Saenko, and Trevor Darrell · 2017
Earlier work this paper cites.
Rouge 2.0: Updated and improved measures for evaluation of summarization tasks
Kavita Ganesan · 2018
Earlier work this paper cites.
Domain generalization with adversarial feature learning
Haoliang Li, Sinno Jialin Pan, Shiqi Wang, and Alex C Kot · 2018
Earlier work this paper cites.
Vqa-e: Explaining, elaborating, and enhancing your answers for visual questions
Qing Li, Qingyi Tao, Shafiq Joty, Jianfei Cai, and Jiebo Luo · 2018
Earlier work this paper cites.
Deep domain generalization via conditional invariant adversarial networks
Ya Li, Xinmei Tian, Mingming Gong, Yajing Liu, Tongliang Liu, Kun Zhang, and Dacheng Tao · 2018
Earlier work this paper cites.
Senti-attend: Image captioning using sentiment and attention
Omid Mohamad Nezami, Mark Dras, Stephen Wan, and Cécile Paris · 2018
Earlier work this paper cites.
Generalizing across domains via cross-gradient training
Shiv Shankar, Vihari Piratla, Soumen Chakrabarti, Siddhartha Chaudhuri, Preethi Jyothi, and Sunita Sarawagi · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
Generalizing to unseen domains via adversarial data augmentation
Riccardo Volpi, Hongseok Namkoong, Ozan Sener, John C Duchi, Vittorio Murino, and Silvio Savarese · 2018
Earlier work this paper cites.
Quanzeng You, Hailin Jin, and Jiebo Luo · 2018
Earlier work this paper cites.
Uniter: Learning universal image-text representations
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2019
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2019
Cited alongside, same era.
Engaging image captioning via personality
Kurt Shuster, Samuel Humeau, Hexiang Hu, Antoine Bordes, and Jason Weston · 2019
Cited alongside, same era.
Unsupervised domain adaptation through self-supervision
Yu Sun, Eric Tzeng, Trevor Darrell, and Alexei A Efros · 2019
Cited alongside, same era.
Visual entailment: A novel task for fine-grained image understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav · 2019
Mesh-Transformer-JAX: Model-Parallel Implementation of Transformer Language Model with JAX
Ben Wang · 2021
Later among the works it cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al · 2021
Later among the works it cites.
Domain generalization with mixstyle
Kaiyang Zhou, Yongxin Yang, Yu Qiao, and Tao Xiang · 2021
Later among the works it cites.
Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
In search of lost domain generalization
Ishaan Gulrajani and David Lopez-Paz · 2020
Cited alongside, same era.
Captioning images taken by people who are blind
Danna Gurari, Yinan Zhao, Meng Zhang, and Nilavra Bhattacharya · 2020
Cited alongside, same era.
Visual news: Benchmark and challenges in news image captioning
Fuxiao Liu, Yinghan Wang, Tianlu Wang, and Vicente Ordonez · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
A survey of unsupervised deep domain adaptation
Garrett Wilson and Diane J Cook · 2020
Cited alongside, same era.
Memcap: Memorizing style knowledge for image captioning
Wentian Zhao, Xinxiao Wu, and Xiaoxun Zhang · 2020
Cited alongside, same era.
Mohamed Afham, Isuru Dissanayake, Dinithi Dissanayake, Amaya Dharmasiri, Kanchana Thilakarathna, and Ranga Rodrigo · 2022
Closest in time.
Clap: Learning audio concepts from natural language supervision
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang · 2022
Closest in time.
Eva: Exploring the limits of masked visual representation learning at scale
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao · 2022
Closest in time.
Audioclip: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel · 2022
Closest in time.
Webly supervised concept expansion for general purpose vision models
Amita Kamath, Christopher Clark, Tanmay Gupta, Eric Kolve, Derek Hoiem, and Aniruddha Kembhavi · 2022
Closest in time.
Simple but effective: Clip embeddings for embodied ai
Apoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2022
Closest in time.
Highmmt: Towards modality and task generalization for high-modality representation learning
Paul Pu Liang, Yiwei Lyu, Xiang Fan, Shengtong Mo, Dani Yogatama, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2022
Closest in time.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Y. Zou · 2022
Closest in time.
Unified-io: A unified model for vision, language, and multi-modal tasks
Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2022
Closest in time.
Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li · 2022
Closest in time.
Text-only training for image captioning using noise-injected clip
David Nukrai, Ron Mokady, and Amir Globerson · 2022
Closest in time.
Clip models are few-shot learners: Empirical studies on vqa and visual entailment
Haoyu Song, Li Dong, Wei-Nan Zhang, Ting Liu, and Furu Wei · 2022
Closest in time.
Language models can see: Plugging visual controls in text generation
Yixuan Su, Tian Lan, Yahui Liu, Fangyu Liu, Dani Yogatama, Yan Wang, Lingpeng Kong, and Nigel Collier · 2022
Closest in time.
Detach and attach: Stylized image captioning without paired stylized dataset
Yutong Tan, Zheng Lin, Peng Fu, Mingyu Zheng, Lanrui Wang, Yanan Cao, and Weipinng Wang · 2022
Closest in time.
Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic
Yoad Tewel, Yoav Shalev, Idan Schwartz, and Lior Wolf · 2022
Closest in time.
Generalizing to unseen domains: A survey on domain generalization
Jindong Wang, Cuiling Lan, Chang Liu, Yidong Ouyang, Tao Qin, Wang Lu, Yiqiang Chen, Wenjun Zeng, and Philip Yu · 2022
Closest in time.
Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang · 2022
Closest in time.
Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, et al · 2022
Closest in time.
Wav2clip: Learning robust audio representations from clip
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello · 2022
Closest in time.
Unified contrastive learning in image-text-label space
Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Bin Xiao, Ce Liu, Lu Yuan, and Jianfeng Gao · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Closest in time.
Multimodal knowledge alignment with reinforcement learning
Youngjae Yu, Jiwan Chung, Heeseung Yun, Jack Hessel, JaeSung Park, Ximing Lu, Prithviraj Ammanabrolu, Rowan Zellers, Ronan Le Bras, Gunhee Kim, et al · 2022
Closest in time.
Socratic models: Composing zero-shot multimodal reasoning with language
Andy Zeng, Maria Attarian, Brian Ichter, Krzysztof Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, and Pete Florence · 2022
Closest in time.
Lit: Zero-shot transfer with locked-image text tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer · 2022
Closest in time.