Fetching the paper…
Reading the bibliography…
Most existing image tokenizers encode images into a fixed number of tokens or patches, overlooking the inherent variability in image complexity.
Improved precision and recall metric for assessing generative models, 2019
Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila · 1904
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning, 2020
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 1911
Earlier work this paper cites.
The jpeg still picture compression standard
G.K. Wallace · 1992
Earlier work this paper cites.
Overview of the h.264/avc video coding standard
T. Wiegand, G.J. Sullivan, G. Bjontegaard, and A. Luthra · 2003
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
Image quality metrics: Psnr vs. ssim
Alain Horé and Djemel Ziou · 2010
Earlier work this paper cites.
Xin Ding, Yongwei Wang, Zuheng Xu, William J. Welch, and Z. Jane Wang · 2011
Earlier work this paper cites.
Auto-encoding variational bayes, 2014
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Generative adversarial networks, 2014
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2015
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation, 2015
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition, 2015
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Improved techniques for training gans, 2016
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Neural discrete representation learning, 2018
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local nash equilibrium, 2018
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2018
Cited alongside, same era.
A style-based generator architecture for generative adversarial networks, 2019
Tero Karras, Samuli Laine, and Timo Aila · 2019
Cited alongside, same era.
Taming transformers for high-resolution image synthesis, 2020
Patrick Esser, Robin Rombach, and Björn Ommer · 2020
Matryoshka representation learning
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, et al · 2022
Later among the works it cites.
Transframer: Arbitrary frame prediction with generative models, 2022
Charlie Nash, João Carreira, Jacob Walker, Iain Barr, Andrew Jaegle, Mateusz Malinowski, and Peter Battaglia · 2022
Later among the works it cites.
Classifier-free diffusion guidance, 2022
Jonathan Ho and Tim Salimans · 2022
Later among the works it cites.
Finite scalar quantization: Vq-vae made simple, 2023
Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen · 2023
Later among the works it cites.
Token merging: Your ViT but faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari · 2020
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
Cited alongside, same era.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh · 2021
Cited alongside, same era.
Automl decathlon: Diverse tasks, modern methods, and efficiency at scale
Nicholas Roberts, Samuel Guo, Cong Xu, Ameet Talwalkar, David Lander, Lvfang Tao, Linhang Cai, Shuaicheng Niu, Jianyu Heng, Hongyang Qin, Minwen Deng, Johannes Hog, Alexander Pfefferle, Sushil Ammanaghatta Shivakumar, Arjun Krishnakumar, Yubo Wang, Rhea Sanjay Sukthanker, Frank Hutter, Euxhen Hasanaj, Tien-Dung Le, Mikhail Khodak, Yuriy Nevmyvaka, Kashif Rasul, Frederic Sala, Anderson Schneider, Junhong Shen, and Evan R. Sparks · 2021
Cited alongside, same era.
Theoretically principled deep rl acceleration via nearest neighbor function approximation
Junhong Shen and Lin F. Yang · 2021
Cited alongside, same era.
Efficient architecture search for diverse tasks
Junhong Shen, Mikhail Khodak, and Ameet Talwalkar · 2022
Cited alongside, same era.
NAS-bench-360: Benchmarking neural architecture search on diverse tasks
Renbo Tu, Nicholas Roberts, Mikhail Khodak, Junhong Shen, Frederic Sala, and Ameet Talwalkar · 2022
Cited alongside, same era.
Later among the works it cites.
Efficient video action detection with token dropout and context refinement, 2023
Lei Chen, Zhan Tong, Yibing Song, Gangshan Wu, and Limin Wang · 2023
Later among the works it cites.
Vision transformers with mixed-resolution tokenization, 2023
Tomer Ronen, Omer Levy, and Avram Golbert · 2023
Later among the works it cites.
Cross-modal fine-tuning: align then refine
Junhong Shen, Liam Li, Lucio M. Dery, Corey Staten, Mikhail Khodak, Graham Neubig, and Ameet Talwalkar · 2023
Later among the works it cites.
Elastictok: Adaptive tokenization for image and video
Wilson Yan, Matei Zaharia, Volodymyr Mnih, Pieter Abbeel, Aleksandra Faust, and Hao Liu · 2024
Later among the works it cites.
Adaptive length image tokenization via recurrent allocation, 2024
Shivam Duggal, Phillip Isola, Antonio Torralba, and William T. Freeman · 2024
Later among the works it cites.
Mu Cai, Jianwei Yang, Jianfeng Gao, and Yong Jae Lee · 2024
Later among the works it cites.
Matryoshka diffusion models, 2024
Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Josh Susskind, and Navdeep Jaitly · 2024
Later among the works it cites.
Matryoshka query transformer for large vision-language models, 2024
Wenbo Hu, Zi-Yi Dou, Liunian Harold Li, Amita Kamath, Nanyun Peng, and Kai-Wei Chang · 2024
Later among the works it cites.
Chameleon: Mixed-modal early-fusion foundation models, 2024
Chameleon Team · 2024
Later among the works it cites.
Transfusion: Predict the next token and diffuse images with one multi-modal model, 2024
Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, and Omer Levy · 2024
Later among the works it cites.