Fetching the paper…
Reading the bibliography…
We propose to replace vector quantization (VQ) in the latent representation of VQ-VAEs with a simple scheme termed finite scalar quantization (FSQ), where we project the VAE representation down to a few dimensions (typically less than 10).
An algorithm for vector quantizer design
Yoseph Linde, Andres Buzo, and Robert Gray · 1980
Earlier work this paper cites.
Vector quantization
Robert Gray · 1984
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
End-to-end optimized image compression
Johannes Ballé, Valero Laparra, and Eero P Simoncelli · 2016
Earlier work this paper cites.
Lossy image compression with compressive autoencoders
Lucas Theis, Wenzhe Shi, Andrew Cunningham, and Ferenc Huszár · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
High quality monocular depth estimation via transfer learning
Ibraheem Alhashim and Peter Wonka · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Conditional probability models for deep image compression
Fabian Mentzer, Eirikur Agustsson, Michael Tschannen, Radu Timofte, and Luc Van Gool · 2018
Earlier work this paper cites.
Joint autoregressive and hierarchical priors for learned image compression
David Minnen, Johannes Ballé, and George D Toderici · 2018
Earlier work this paper cites.
Theory and experiments on vector quantized autoencoders
Aurko Roy, Ashish Vaswani, Arvind Neelakantan, and Niki Parmar · 2018
Earlier work this paper cites.
Assessing generative models via precision and recall
Mehdi SM Sajjadi, Olivier Bachem, Mario Lucic, Olivier Bousquet, and Sylvain Gelly · 2018
Earlier work this paper cites.
Deep generative models for distribution-preserving lossy compression
Michael Tschannen, Eirikur Agustsson, and Mario Lucic · 2018
Earlier work this paper cites.
Generative adversarial networks for extreme learned image compression
Eirikur Agustsson, Michael Tschannen, Fabian Mentzer, Radu Timofte, and Luc Van Gool · 2019
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli · 2019
Earlier work this paper cites.
Piano genie
Chris Donahue, Ian Simon, and Sander Dieleman · 2019
Earlier work this paper cites.
Dvc: An end-to-end deep video compression framework
Guo Lu, Wanli Ouyang, Dong Xu, Xiaoyun Zhang, Chunlei Cai, and Zhiyong Gao · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Differentiable product quantization for end-to-end embedding compression
Ting Chen, Lala Li, and Yizhou Sun · 2020
Cited alongside, same era.
Learned image compression with discretized gaussian mixture likelihoods and attention modules
Zhengxue Cheng, Heming Sun, Masaru Takeuchi, and Jiro Katto · 2020
Cited alongside, same era.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2020
Cited alongside, same era.
Maskgit: Masked generative image transformer
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman · 2022
Later among the works it cites.
Image compression with product quantized masked image modeling
Alaaeldin El-Nouby, Matthew J Muckley, Karen Ullrich, Ivan Laptev, Jakob Verbeek, and Hervé Jégou · 2022
Later among the works it cites.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2022
Later among the works it cites.
Uvim: A unified modeling approach for vision with learned guiding codes
Alexander Kolesnikov, André Susano Pinto, Lucas Beyer, Xiaohua Zhai, Jeremiah Harmsen, and Neil Houlsby · 2022
Later among the works it cites.
Autoregressive image generation using residual quantization
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taming transformers for high-resolution image synthesis. 2021 ieee
Patrick Esser, Robin Rombach, and Björn Ommer · 2020
Cited alongside, same era.
Robust training of vector quantized bottleneck models
Adrian Łańcucki, Jan Chorowski, Guillaume Sanchez, Ricard Marxer, Nanxin Chen, Hans JGA Dolfing, Sameer Khurana, Tanel Alumäe, and Antoine Laurent · 2020
Cited alongside, same era.
High-fidelity generative image compression
Fabian Mentzer, George D Toderici, Michael Tschannen, and Eirikur Agustsson · 2020
Cited alongside, same era.
Hierarchical quantized autoencoders
Will Williams, Sam Ringer, Tom Ash, David MacLeod, Jamie Dougherty, and John Hughes · 2020
Cited alongside, same era.
Beit: Bert pre-training of image transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Cited alongside, same era.
Variable-rate discrete representation learning
Sander Dieleman, Charlie Nash, Jesse Engel, and Karen Simonyan · 2021
Cited alongside, same era.
Yuhta Takida, Takashi Shibuya, WeiHsiang Liao, Chieh-Hsin Lai, Junki Ohmura, Toshimitsu Uesaka, Naoki Murata, Shusuke Takahashi, Toshiyuki Kumakura, and Yuki Mitsufuji · 2022
Later among the works it cites.
Phenaki: Variable length video generation from open domain textual descriptions
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan · 2022
Later among the works it cites.
Scaling laws for generative mixed-modal language models
Armen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu, Karen Hambardzumyan, Susan Zhang, Stephen Roller, Naman Goyal, Omer Levy, and Luke Zettlemoyer · 2023
Closest in time.
Muse: Text-to-image generation via masked generative transformers
Huiwen Chang, Han Zhang, Jarred Barber, AJ Maschinot, Jose Lezama, Lu Jiang, Ming-Hsuan Yang, Kevin Murphy, William T Freeman, Michael Rubinstein, et al · 2023
Closest in time.
ADM TensorFlow Suite, 2023
Prafulla Dhariwal and Alexander Nichol · 2023
Closest in time.
Disentanglement via latent quantization
Kyle Hsu, Will Dorrell, James CR Whittington, Jiajun Wu, and Chelsea Finn · 2023
Closest in time.
Not all image regions matter: Masked vector quantization for autoregressive image generation
Mengqi Huang, Zhendong Mao, Quan Wang, and Yongdong Zhang · 2023
Closest in time.
Minyoung Huh, Brian Cheung, Pulkit Agrawal, and Phillip Isola · 2023
Closest in time.
Magvlt: Masked generative vision-and-language transformer
Sungwoong Kim, Daejin Jo, Donghoon Lee, and Jongmin Kim · 2023
Closest in time.
Mage: Masked generative encoder to unify representation learning and image synthesis
Tianhong Li, Huiwen Chang, Shlok Mishra, Han Zhang, Dina Katabi, and Dilip Krishnan · 2023
Closest in time.
M2t: Masking transformers twice for faster decoding
Fabian Mentzer, Eirikur Agustsson, and Michael Tschannen · 2023
Closest in time.