Fetching the paper…
Reading the bibliography…
Multimodal image-text models have shown remarkable performance in the past few years.
Visual entailment: A novel task for fine-grained image understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav · 1901
Earlier work this paper cites.
Augmix: A simple data processing method to improve robustness and uncertainty
Dan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan · 1912
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Bert-attack: Adversarial attack against bert using bert
Linyang Li, Ruotian Ma, Qipeng Guo, X. Xue, and Xipeng Qiu · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan Conrad Bovik, Hamid R. Sheikh, and Eero P. Simoncelli · 2004
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, A. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael J. Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Generating images from captions with attention
Elman Mansimov, Emilio Parisotto, Jimmy Ba, and Ruslan Salakhutdinov · 2016
Earlier work this paper cites.
Improving the robustness of deep neural networks via stability training
Stephan Zheng, Yang Song, Thomas Leung, and Ian J. Goodfellow · 2016
Earlier work this paper cites.
Boosting adversarial attacks with momentum
Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization
Xun Huang and Serge J. Belongie · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R. Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
A corpus of natural language for visual reasoning
Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi · 2017
Earlier work this paper cites.
Fooling vision and language models despite localization and attention mechanism
Xiaojun Xu, Xinyun Chen, Chang Liu, Anna Rohrbach, Trevor Darrell, and Dawn Xiaodong Song · 2017
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk · 2018
Earlier work this paper cites.
On adversarial examples for character-level neural machine translation
J. Ebrahimi, Daniel Lowd, and Dejing Dou · 2018
Earlier work this paper cites.
Delete, retrieve, generate: a simple approach to sentiment and style transfer
Juncen Li, Robin Jia, He He, and Percy Liang · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman · 2018
Earlier work this paper cites.
Visual entailment task for visually-grounded language learning
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
Formality style transfer for noisy, user-generated conversations: Extracting labeled, parallel data from unlabeled corpora
Isak Czeresnia Etinger and Alan W. Black · 2019
Earlier work this paper cites.
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix Wichmann, and Wieland Brendel · 2019
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas G. Dietterich · 2019
Earlier work this paper cites.
Are you looking? grounding to multiple modalities in vision-and-language navigation
Ronghang Hu, Daniel Fried, Anna Rohrbach, Dan Klein, Trevor Darrell, and Kate Saenko · 2019
Earlier work this paper cites.
Nesterov accelerated gradient and scale invariance for adversarial attacks
Jiadong Lin, Chuanbiao Song, Kun He, Liwei Wang, and John E. Hopcroft · 2019
Earlier work this paper cites.
Nlp augmentation
Edward Ma · 2019
Earlier work this paper cites.
Benchmarking robustness in object detection: Autonomous driving when winter is coming
Claudio Michaelis et al · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Models in the wild: On corruption robustness of neural nlp systems
Barbara Rychalska, Dominika Basaj, Alicja Gosiewska, and P. Biecek · 2019
Earlier work this paper cites.
Eda: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou · 2019
Earlier work this paper cites.
Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations
Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang, Lei Li, Weiwei Sun, and Wei-Ying Ma · 2019
Earlier work this paper cites.
A fourier perspective on model robustness in computer vision
Dong Yin, Raphael Gontijo Lopes, Jonathon Shlens, Ekin Dogus Cubuk, and Justin Gilmer · 2019
Earlier work this paper cites.
Experience grounds language
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Yue Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph P. Turian · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown et al · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Robustbench: a standardized adversarial robustness benchmark
Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein · 2020
Cited alongside, same era.
Large-scale adversarial training for vision-and-language representation learning
Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu · 2020
Pali: A jointly-scaled multilingual language-image model
Xi Chen et al · 2022
Closest in time.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generative transformers
Jaemin Cho, Abhaysinh Zala, and Mohit Bansal · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery et al · 2022
Closest in time.
Discovering the hidden vocabulary of dalle-2
Giannis Daras and Alexandros G Dimakis · 2022
Closest in time.
Multimodal automl for image, text and tabular data
Nick Erickson, Xingjian Shi, James Sharpnack, and Alexander J. Smola · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Facebook fair’s wmt19 news translation task submission
Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov · 2020
Cited alongside, same era.
Generative text style transfer for improved language sophistication
Robert Schmidt · 2020
Cited alongside, same era.
Measuring robustness to natural distribution shifts in image classification
Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt · 2020
Cited alongside, same era.
Cat-gen: Improving robustness in nlp models via controlled adversarial text generation
Tianlu Wang, Xuezhi Wang, Yao Qin, Ben Packer, Kang Li, Jilin Chen, Alex Beutel, and Ed H. Chi · 2020
Cited alongside, same era.
Reveal of vision transformers robustness against adversarial attacks
Ahmed Aldahdooh, Wassim Hamidouche, and Olivier Déforges · 2021
Cited alongside, same era.
Understanding robustness of transformers for image classification
Srinadh Bhojanapalli et al · 2021
Cited alongside, same era.
Robustness and adversarial examples in natural language processing
Kai-Wei Chang, He He, Robin Jia, and Sameer Singh · 2021
Cited alongside, same era.
Data determines distributional robustness in contrastive language image pre-training (clip)
Alexander W. Fang, Gabriel Ilharco, Mitchell Wortsman, Yuhao Wan, Vaishaal Shankar, Achal Dave, and Ludwig Schmidt · 2022
Closest in time.
Vision models are more robust and fair when pretrained on uncurated images without supervision
Priya Goyal, Quentin Duval, Isaac Seessel, Mathilde Caron, Ishan Misra, Levent Sagun, Armand Joulin, and Piotr Bojanowski · 2022
Closest in time.
Grit: General robust image task benchmark
Tanmay Gupta, Ryan Marten, Aniruddha Kembhavi, and Derek Hoiem · 2022
Closest in time.
An empirical exploration of cross-domain alignment between language and electroencephalogram
William Han, Jielin Qiu, Jiacheng Zhu, Mengdi Xu, Douglas Weber, Bo Li, and Ding Zhao · 2022
Closest in time.
Mixgen: A new multi-modal data augmentation
Xiaoshuai Hao, Yi Zhu, Srikar Appalaraju, Aston Zhang, Wanqian Zhang, Boyang Li, and Mu Li · 2022
Closest in time.
Scaling up vision-language pretraining for image captioning
Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang · 2022
Closest in time.
Transformers are adaptable task planners
Vidhi Jain, Yixin Lin, Eric Undersander, Yonatan Bisk, and Akshara Rai · 2022
Closest in time.
The role of imagenet classes in fréchet inception distance
Tuomas Kynkaanniemi, Tero Karras, Miika Aittala, Timo Aila, and Jaakko Lehtinen · 2022
Closest in time.
Foundations and recent trends in multimodal machine learning: Principles, challenges, and open questions
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency · 2022
Closest in time.
Design guidelines for prompt engineering text-to-image generative models
Vivian Liu and Lydia B. Chilton · 2022
Closest in time.
Unified-io: A unified model for vision, language, and multi-modal tasks
Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2022
Closest in time.
The king is naked: on the notion of robustness for natural language processing
Emanuele La Malfa and Marta Z. Kwiatkowska · 2022
Closest in time.
Film: Following instructions in language with modular methods
So Yeon Min, Devendra Singh Chaplot, Pradeep Ravikumar, Yonatan Bisk, and Ruslan Salakhutdinov · 2022
Closest in time.
Grit: Faster and better image captioning transformer using dual visual features
Van-Quang Nguyen, Masanori Suganuma, and Takayuki Okatani · 2022
Closest in time.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen · 2022
Closest in time.
On aliased resizing and surprising subtleties in gan evaluation
Gaurav Parmar, Richard Zhang, and Jun-Yan Zhu · 2022
Closest in time.
Perception test : A diagnostic benchmark for multimodal models
Viorica Patraucean et al · 2022
Closest in time.
Vision transformers are robust learners
Sayak Paul and Pin-Yu Chen · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
Robin Rombach, A. Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Closest in time.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia et al · 2022
Closest in time.
Multi-modal robustness analysis against language and visual perturbations
Madeline Chantry Schiappa, Yogesh Singh Rawat, Shruti Vyas, Vibhav Vineet, and Hamid Palangi · 2022
Closest in time.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann et al · 2022
Closest in time.
Assaying out-of-distribution generalization in transfer learning
F. Wenzel et al · 2022
Closest in time.
Vision-language pre-training with triple contrastive learning
Jinyu Yang, Jiali Duan, S. Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul M. Chilimbi, and Junzhou Huang · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Closest in time.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Closest in time.
Towards adversarial attack on vision-language pre-training models
Jiaming Zhang, Qiaomin Yi, and Jitao Sang · 2022
Closest in time.
Understanding the robustness in vision transformers
Daquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao, Anima Anandkumar, Jiashi Feng, and José Manuel Álvarez · 2022
Closest in time.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi · 2023
Closest in time.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Closest in time.