Fetching the paper…
Reading the bibliography…
Visual autoregressive models typically adhere to a raster-order ``next-token prediction" paradigm, which overlooks the spatial and temporal locality inherent in visual content.
A 3x3 Isotropic Gradient Operator for Image Processing , pages 271–272
I. Sobel and G. Feldman · 1973
Earlier work this paper cites.
A computational approach to edge detection
John Canny · 1986
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
The jpeg still picture compression standard
Gregory K Wallace · 1991
Earlier work this paper cites.
Png (portable network graphics) specification version 1.0
Thomas Boutell · 1997
Earlier work this paper cites.
Bilateral filtering for gray and color images
Carlo Tomasi and Roberto Manduchi · 1998
Earlier work this paper cites.
Image denoising by sparse 3-d transform-domain collaborative filtering
Kostadin Dabov, Alessandro Foi, Vladimir Katkovnik, and Karen Egiazarian · 2007
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K Soomro · 2012
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation. arxiv 2015
Jonathan Long, Evan Shelhamer, Trevor Darrell, and UC Berkeley · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan · 2014
Earlier work this paper cites.
Deep residual learning for image recognition. arxiv e-prints
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
You only look once: unified, real-time object detection (2015)
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2015
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al · 2016
Earlier work this paper cites.
Pixelsnail: An improved autoregressive generative model, 2017
Xi Chen, Nikhil Mishra, Mostafa Rohaninejad, and Pieter Abbeel · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Attention is all you need
A Waswani, N Shazeer, N Parmar, J Uszkoreit, L Jones, A Gomez, L Kaiser, and I Polosukhin · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford · 2018
Cited alongside, same era.
Painting outside the box: Image outpainting with gans
Mark Sabini and Gili Rusak · 2018
Cited alongside, same era.
Towards accurate generative models of video: A new metric & challenges
Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Attention augmented convolutional networks, 2020
Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V. Le · 2020
Cited alongside, same era.
Videofusion: Decomposed diffusion models for high-quality video generation
Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan · 2023
Later among the works it cites.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach · 2023
Later among the works it cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
End-to-end object detection with transformers, 2020
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy · 2020
Cited alongside, same era.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al · 2020
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Cited alongside, same era.
Edge guided progressively generative image outpainting
Han Lin, Maurice Pagnucco, and Yang Song · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows, 2021
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
Haoge Deng, Ting Pan, Haiwen Diao, Zhengxiong Luo, Yufeng Cui, Huchuan Lu, Shiguang Shan, Yonggang Qi, and Xinlong Wang · 2024
Later among the works it cites.
Geneval: An object-focused framework for evaluating text-to-image alignment
Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt · 2024
Later among the works it cites.
Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis
Jian Han, Jinlai Liu, Yi Jiang, Bin Yan, Yuqi Zhang, Zehuan Yuan, Bingyue Peng, and Xiaobing Liu · 2024
Later among the works it cites.
Zipar: Accelerating autoregressive image generation through spatial locality
Yefei He, Feng Chen, Yuanyu He, Shaoxuan He, Hong Zhou, Kaipeng Zhang, and Bohan Zhuang · 2024
Later among the works it cites.
Dongyang Liu, Shitian Zhao, Le Zhuo, Weifeng Lin, Yu Qiao, Hongsheng Li, and Peng Gao · 2024
Later among the works it cites.
Jinliang Lu, Ziliang Pang, Min Xiao, Yaochen Zhu, Rui Xia, and Jiajun Zhang · 2024
Later among the works it cites.
Hierarchical patch diffusion models for high-resolution video generation
Ivan Skorokhodov, Willi Menapace, Aliaksandr Siarohin, and Sergey Tulyakov · 2024
Later among the works it cites.
Autoregressive model beats diffusion: Llama for scalable image generation
Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan · 2024
Later among the works it cites.
Chameleon: Mixed-modal early-fusion foundation models
Chameleon Team · 2024
Later among the works it cites.
Accelerating auto-regressive text-to-image generation with training-free speculative jacobi decoding
Yao Teng, Han Shi, Xian Liu, Xuefei Ning, Guohao Dai, Yu Wang, Zhenguo Li, and Xihui Liu · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang · 2024
Later among the works it cites.
Janus-pro: Unified multimodal understanding and generation with data and model scaling
Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, and Chong Ruan · 2025
Closest in time.