Fetching the paper…
Reading the bibliography…
While the integration of transformers in vision models have yielded significant improvements on vision tasks they still require significant amounts of computation for both training and inference.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol · 2010
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation, 2015
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop
Fisher Yu, Yinda Zhang, Shuran Song, Ari Seff, and Jianxiong Xiao · 2015
Earlier work this paper cites.
A note on the evaluation of generative models, 2016
Lucas Theis, Aäron van den Oord, and Matthias Bethge · 2016
Earlier work this paper cites.
A learned representation for artistic style
Vincent Dumoulin, Jonathon Shlens, and Manjunath Kudlur · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization, 2017
Xun Huang and Serge Belongie · 2017
Earlier work this paper cites.
Demystifying MMD GANs
Mikołaj Bińkowski, Dougal J. Sutherland, Michael Arbel, and Arthur Gretton · 2018
Earlier work this paper cites.
Progressive growing of GANs for improved quality, stability, and variation
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen · 2018
Earlier work this paper cites.
Which training methods for GANs do actually converge?
Lars Mescheder, Andreas Geiger, and Sebastian Nowozin · 2018
Earlier work this paper cites.
Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel · 2019
Earlier work this paper cites.
Ganalyze: Toward visual definitions of cognitive image properties
Lore Goetschalckx, Alex Andonian, Aude Oliva, and Phillip Isola · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila · 2019
Earlier work this paper cites.
Improved precision and recall metric for assessing generative models
Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila · 2019
Earlier work this paper cites.
Stand-alone self-attention in vision models
Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens · 2019
Earlier work this paper cites.
Self-attention generative adversarial networks
Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena · 2019
Earlier work this paper cites.
Hype: A benchmark for human eye perceptual evaluation of generative models
Sharon Zhou, Mitchell Gordon, Ranjay Krishna, Austin Narcomey, Li F Fei-Fei, and Michael Bernstein · 2019
Earlier work this paper cites.
Longformer: The long-document transformer, 2020
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2020
Earlier work this paper cites.
The origins and prevalence of texture bias in convolutional neural networks
Katherine Hermann, Ting Chen, and Simon Kornblith · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Jsi-gan: Gan-based joint super-resolution and inverse tone-mapping with pixel-wise task-specific filters for uhd hdr video
Soo Ye Kim, Jihyong Oh, and Munchurl Kim · 2020
Earlier work this paper cites.
Reliable fidelity and diversity metrics for generative models
Muhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi, and Jaejun Yoo · 2020
Earlier work this paper cites.
Neural supersampling for real-time rendering
Lei Xiao, Salah Nouri, Matt Chapman, Alexander Fix, Douglas Lanman, and Anton Kaplanyan · 2020
Earlier work this paper cites.
Approximation capabilities of neural ODEs and invertible residual networks
Han Zhang, Xi Gao, Jacob Unterman, and Tom Arodz · 2020
Cited alongside, same era.
Mobilestylegan: A lightweight convolutional neural network for high-fidelity image synthesis, 2021
Sergei Belousov · 2021
Cited alongside, same era.
Investigating softmax tempering for training neural machine translation models
Raj Dabre and Atsushi Fujita · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Interactive path tracing and reconstruction of sparse volumes
Nikolai Hofmann, Jon Hasselgren, Petrik Clarberg, and Jacob Munkberg · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Closest in time.
Stylegan-xl: Scaling stylegan to large diverse datasets
Axel Sauer, Katja Schwarz, and Andreas Geiger · 2022
Closest in time.
How to train your vit? data, augmentation, and regularization in vision transformers
Andreas Peter Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer · 2022
Closest in time.
Generative flows with invertible attentions
Rhea Sanjay Sukthanker, Zhiwu Huang, Suryansh Kumar, Radu Timofte, and Luc Van Gool · 2022
Closest in time.
A tri-layer plugin to improve occluded detection
Guanqi Zhan, Weidi Xie, and Andrew Zisserman · 2022
Closest in time.
Styleswin: Transformer-based gan for high-resolution image generation
Bowen Zhang, Shuyang Gu, Bo Zhang, Jianmin Bao, Dong Chen, Fang Wen, Yong Wang, and Baining Guo · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Drew A Hudson and C. Lawrence Zitnick · 2021
Cited alongside, same era.
Compositional transformers for scene generation
Drew A Hudson and C. Lawrence Zitnick · 2021
Cited alongside, same era.
Alias-free generative adversarial networks
Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
On self-supervised image representations for gan evaluation
Stanislav Morozov, Andrey Voynov, and Artem Babenko · 2021
Cited alongside, same era.
Real-time neural radiance caching for path tracing
Thomas Müller, Fabrice Rousselle, Jan Novák, and Alexander Keller · 2021
Cited alongside, same era.
Projected gans converge faster
Axel Sauer, Kashyap Chitta, Jens Müller, and Andreas Geiger · 2021
Cited alongside, same era.
Closest in time.
Improving image generation with better captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al · 2023
Closest in time.
Swin-unet: Unet-like pure transformer for medical image segmentation
Hu Cao, Yueyue Wang, Joy Chen, Dongsheng Jiang, Xiaopeng Zhang, Qi Tian, and Manning Wang · 2023
Closest in time.
Quan Dao, Hao Phung, Binh Nguyen, and Anh Tran · 2023
Closest in time.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Closest in time.
Neighborhood attention transformer
Ali Hassani, Steven Walton, Jiachen Li, Shen Li, and Humphrey Shi · 2023
Closest in time.
Oneformer: One transformer to rule universal image segmentation
Jitesh Jain, Jiachen Li, Mang Tik Chiu, Ali Hassani, Nikita Orlov, and Humphrey Shi · 2023
Closest in time.
The role of imagenet classes in fréchet inception distance
Tuomas Kynkäänniemi, Tero Karras, Miika Aittala, Timo Aila, and Jaakko Lehtinen · 2023
Closest in time.
Compensation sampling for improved convergence in diffusion models, 2023
Hui Lu, Albert ali Salah, and Ronald Poppe · 2023
Closest in time.
Are emergent abilities of large language models a mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2023
Closest in time.
Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models
George Stein, Jesse C. Cresswell, Rasa Hosseinzadeh, Yi Sui, Brendan Leigh Ross, Valentin Villecroze, Zhaoyan Liu, Anthony L. Caterini, Eric Taylor, and Gabriel Loaiza-Ganem · 2023
Closest in time.
Diffusion-GAN: Training GANs with diffusion
Zhendong Wang, Huangjie Zheng, Pengcheng He, Weizhu Chen, and Mingyuan Zhou · 2023
Closest in time.
Scalable high-resolution pixel-space image synthesis with hourglass diffusion transformers
Katherine Crowson, Stefan Andreas Baumann, Alex Birch, Tanishq Mathew Abraham, Daniel Z Kaplan, and Enrico Shippole · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis, 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach · 2024
Closest in time.
Faster neighborhood attention: Reducing the o(n2̂) cost of self attention at the threadblock level, 2024
Ali Hassani, Wen-Mei Hwu, and Humphrey Shi · 2024
Closest in time.
SDXL: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach · 2024
Closest in time.
Fast high-resolution image synthesis with latent adversarial diffusion distillation
Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and Robin Rombach · 2024
Closest in time.
Is Less More? Rendering for Esports
Benjamin Watson, Josef Spjut, Joohwan Kim, Byungjoo Lee, Mijin Yoo, Peter Shirley, and Rulon Raymond · 2024
Closest in time.