Fetching the paper…
Reading the bibliography…
This work examines the challenges of training neural networks using vector quantization using straight-through estimation.
Vector quantization
Gray, R · 1984
Earlier work this paper cites.
Improved versions of learning vector quantization
Kohonen, T · 1990
Earlier work this paper cites.
Clustering with bregman divergences
Banerjee, A., Merugu, S., Dhillon, I. S., Ghosh, J., and Lafferty, J · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Liu, Z., Luo, P., Wang, X., and Tang, X · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Kullback-leibler divergence and moment matching for hyperspherical probability distributions
Kurz, G., Pfaff, F., and Hanebeck, U. D · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2017
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Continuous relaxation training of discrete latent variable image models
Sønderby, C. K., Poole, B., and Mnih, A · 2017
Cited alongside, same era.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Cited alongside, same era.
Deep clustering for unsupervised learning of visual features
Caron, M., Bojanowski, P., Joulin, A., and Douze, M · 2018
Cited alongside, same era.
Kl divergence for machine learning
Ghosh, D · 2018
Cited alongside, same era.
Fast decoding in sequence models using discrete latent variables
Kaiser, L., Bengio, S., Roy, A., Vaswani, A., Parmar, N., Uszkoreit, J., and Shazeer, N · 2018
Cited alongside, same era.
Theory and experiments on vector quantized autoencoders
Roy, A., Vaswani, A., Neelakantan, A., and Parmar, N · 2018
Contrastive multiview coding
Tian, Y., Krishnan, D., and Isola, P · 2020
Later among the works it cites.
Hierarchical quantized autoencoders
Williams, W., Ringer, S., Ash, T., MacLeod, D., Dougherty, J., and Hughes, J · 2020
Later among the works it cites.
deep-vector-quantization
Karpathy, A · 2021
Later among the works it cites.
Vector quantized models for planning
Ozair, S., Li, Y., Razavi, A., Antonoglou, I., Van Den Oord, A., and Vinyals, O · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Later among the works it cites.
Videogpt: Video generation using vq-vae and transformers
Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
Self-labelling via simultaneous clustering and representation learning
Asano, Y. M., Rupprecht, C., and Vedaldi, A · 2020
Cited alongside, same era.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A · 2020
Cited alongside, same era.
Vector-quantized autoregressive predictive coding
Chung, Y.-A., Tang, H., and Glass, J · 2020
Cited alongside, same era.
Jukebox: A generative model for music
Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I · 2020
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2021
Later among the works it cites.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Later among the works it cites.
When vision transformers outperform resnets without pre-training or strong data augmentations
Chen, X., Hsieh, C.-J., and Gong, B · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Later among the works it cites.
Autoregressive image generation using residual quantization
Lee, D., Kim, C., Kim, S., Cho, M., and Han, W.-S · 2022
Later among the works it cites.
Sq-vae: Variational bayes on discrete representation with self-annealed stochastic quantization
Takida, Y., Shibuya, T., Liao, W., Lai, C.-H., Ohmura, J., Uesaka, T., Murata, N., Takahashi, S., Kumakura, T., and Mitsufuji, Y · 2022
Later among the works it cites.
Vector-quantized image modeling with improved vqgan
Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y · 2022
Later among the works it cites.
Muse: Text-to-image generation via masked generative transformers
Chang, H., Zhang, H., Barber, J., Maschinot, A., Lezama, J., Jiang, L., Yang, M.-H., Murphy, K., Freeman, W. T., Rubinstein, M., et al · 2023
Closest in time.