Fetching the paper…
Reading the bibliography…
Mirroring the success of masked language models, vision-and-language counterparts like ViLBERT, LXMERT and UNITER have achieved state of the art performance on a variety of multimodal discriminative tasks like visual question answering and visual grounding.
ImageNet: A Large-Scale Hierarchical Image Database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Generative Adversarial Networks
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Gaussian Error Linear Units (GELUs)
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
Perceptual Losses for Real-Time Style Transfer and Super-Resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li Jia-Li, David Ayman Shamma, Michael Bernstein, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles
Mehdi Noroozi and Paolo Favaro. 2016 · 2016
Earlier work this paper cites.
Pixel Recurrent Neural Networks
Aaron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016 · 2016
Earlier work this paper cites.
Improved Techniques for Training GANs
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016 · 2016
Earlier work this paper cites.
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2016 · 2016
Earlier work this paper cites.
Instance Normalization: The Missing Ingredient for Fast Stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. 2016 · 2016
Earlier work this paper cites.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Earlier work this paper cites.
Visual7W: Grounded Question Answering in Images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization
Xun Huang and Serge Belongie. 2017 · 2017
Earlier work this paper cites.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2017 · 2017
Earlier work this paper cites.
Jae Hyun Lim and Jong Chul Ye. 2017 · 2017
Earlier work this paper cites.
Plug & Play Generative Networks: Conditional Iterative Generation of Images in Latent Space
Anh Nguyen, Jeff Clune, Yoshua Bengio, Alexey Dosovitskiy, and Jason Yosinski. 2017 · 2017
Earlier work this paper cites.
Conditional Image Synthesis With Auxiliary Classifier GANs
Augustus Odena, Christopher Olah, and Jonathon Shlens. 2017 · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chana, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017 · 2017
Cited alongside, same era.
Hierarchical implicit models and likelihood-free variational inference
Dustin Tran, Rajesh Ranganath, and David M. Blei. 2017 · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
StackGAN : Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris Metaxas. 2017 · 2017
A Generalized Framework of Sequence Generation with Application to Undirected Sequence Models
Elman Mansimov, Alex Wang, and Kyunghyun Cho. 2019 · 2019
Later among the works it cites.
Semantic Image Synthesis with Spatially-Adaptive Normalization
Taesung Park, Ming-yu Liu, Ting-chun Wang, and Jun-yan Zhu. 2019 · 2019
Later among the works it cites.
Generating Diverse High-Fidelity Images with VQ-VAE-2
Ali Razavi, Aaron van den Oord, and Oriol Vinyals. 2019 · 2019
Later among the works it cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2019 · 2019
Later among the works it cites.
A Corpus for Reasoning About Natural Language Grounded in Photographs
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Spectral Normalization for Generative Adversarial Networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. 2018 · 2018
Cited alongside, same era.
cGANs with Projection Discriminator
Takeru Miyato and Masanori Koyama. 2018 · 2018
Cited alongside, same era.
Representation Learning with Contrastive Predictive Coding
Aaron Van Den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Cited alongside, same era.
High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs
Ting-chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. 2018 · 2018
Cited alongside, same era.
AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. 2018 · 2018
Cited alongside, same era.
StackGAN++ : Realistic Image Synthesis with Stacked Generative Adversarial Networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, and Senior Member. 2018 · 2018
Cited alongside, same era.
Fusion of Detected Objects in Text for Visual Question Answering
Chris Alberti, Jeffrey Ling, Michael Collins, and David Reitter. 2019 · 2019
Cited alongside, same era.
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019 · 2019
Later among the works it cites.
VideoBERT: A Joint Model for Video and Language Representation Learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019 · 2019
Later among the works it cites.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
Selfie: Self-supervised Pretraining for Image Embedding
Trieu H. Trinh, Minh-Thang Luong, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model
Alex Wang and Kyunghyun Cho. 2019 · 2019
Later among the works it cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 2019
Later among the works it cites.
DM-GAN: Dynamic memory generative adversarial networks for text-to-image synthesis
Minfeng Zhu, Pingbo Pan, Wei Chen, and Yi Yang. 2019 · 2019
Later among the works it cites.
UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Songhao Piao, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2020 · 2020
Closest in time.
Learning Representations by Predicting Bags of Visual Words
Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, and Matthieu Cord. 2020 · 2020
Closest in time.
Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. 2020 · 2020
Closest in time.
Probing Text Models for Common Ground with Visual Representations
Gabriel Ilharco, Rowan Zellers, Ali Farhadi, and Hannaneh Hajishirzi. 2020 · 2020
Closest in time.
In Defense of Grid Features for Visual Question Answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik Learned-Miller, and Xinlei Chen. 2020 · 2020
Closest in time.
Analyzing and Improving the Image Quality of StyleGAN
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020 · 2020
Closest in time.
Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training
Gen Li, Nan Duan, Yuejian Fang, Daxin Jiang, and Ming Zhou. 2020 · 2020
Closest in time.
Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word Order
Yi Liao, Xin Jiang, and Qun Liu. 2020 · 2020
Closest in time.
12-in-1: Multi-Task Vision and Language Representation Learning
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee. 2020 · 2020
Closest in time.
ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data
Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sachet. 2020 · 2020
Closest in time.
M-BERT: Injecting Multimodal Information in the BERT Structure
Wasifur Rahman, Md Kamrul Hasan, Amir Zadeh, Louis-Philippe Morency, and Mohammed Ehsan Hoque. 2020 · 2020
Closest in time.
Unified Vision-Language Pre-Training for Image Captioning and VQA
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. 2020 · 2020
Closest in time.