Fetching the paper…
Reading the bibliography…
Current vision systems typically assign fixed-length representations to images, regardless of the information content.
On the measure of intelligence, 2019
François Chollet · 1911
Earlier work this paper cites.
Low-complexity art
Jürgen Schmidhuber · 1996
Earlier work this paper cites.
Tongzhou Wang and Phillip Isola · 2005
Earlier work this paper cites.
The hutter prize. http://prize.hutter1.net, 2006
Marcus Hutter · 2006
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
Shane Legg and Marcus Hutter · 2007
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Pushmeet Kohli Nathan Silberman, Derek Hoiem and Rob Fergus · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Earlier work this paper cites.
Detecting people in artwork with cnns
Nicholas Westlake, Hongping Cai, and Peter Hall · 2016
Earlier work this paper cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2018
Earlier work this paper cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2018
Earlier work this paper cites.
Savoias: A diverse, multi-category visual complexity dataset
Elham Saraee, Mona Jalal, and Margrit Betke · 2018
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis, 2020
Patrick Esser, Robin Rombach, and Björn Ommer · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Cited alongside, same era.
An empirical study of training self-supervised vision transformers
Xinlei Chen, Saining Xie, and Kaiming He · 2021
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2021
Cited alongside, same era.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh · 2021
Cited alongside, same era.
Token merging: Your ViT but faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman · 2023
Later among the works it cites.
Minyoung Huh, Brian Cheung, Pulkit Agrawal, and Phillip Isola · 2023
Later among the works it cites.
Scalable adaptive computation for iterative generation, 2023
Allan Jabri, David Fleet, and Ting Chen · 2023
Later among the works it cites.
Adaptive computation with elastic input sequence, 2023
Fuzhao Xue, Valerii Likhosherstov, Anurag Arnab, Neil Houlsby, Mostafa Dehghani, and Yang You · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks, 2021
Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang, Uzi Vishkin, Micah Goldblum, and Tom Goldstein · 2021
Cited alongside, same era.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork · 2021
Cited alongside, same era.
Vector-quantized image modeling with improved VQGAN
Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu · 2021
Cited alongside, same era.
Flexivit: One model for all patch sizes
Lucas Beyer, Pavel Izmailov, Alexander Kolesnikov, Mathilde Caron, Simon Kornblith, Xiaohua Zhai, Matthias Minderer, Michael Tschannen, Ibrahim Alabdulmohsin, and Filip Pavetic · 2022
Cited alongside, same era.
Scalable adaptive computation for iterative generation, 2022
Allan Jabri, David Fleet, and Ting Chen · 2022
Cited alongside, same era.
Auto-encoding variational bayes, 2022
Diederik P Kingma and Max Welling · 2022
Cited alongside, same era.
Matryoshka representation learning
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, et al · 2022
Cited alongside, same era.
Mu Cai, Jianwei Yang, Jianfeng Gao, and Yong Jae Lee · 2024
Closest in time.
Think before you speak: Training language models with pause tokens, 2024
Sachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar, and Vaishnavh Nagarajan · 2024
Closest in time.
Thinking tokens for language modeling, 2024
David Herel and Tomas Mikolov · 2024
Closest in time.
Matryoshka query transformer for large vision-language models, 2024
Wenbo Hu, Zi-Yi Dou, Liunian Harold Li, Amita Kamath, Nanyun Peng, and Kai-Wei Chang · 2024
Closest in time.
Mixture of nested experts: Adaptive processing of visual tokens, 2024
Gagan Jain, Nidhi Hegde, Aditya Kusupati, Arsha Nagrani, Shyamal Buch, Prateek Jain, Anurag Arnab, and Sujoy Paul · 2024
Closest in time.
Elastictok: Adaptive tokenization for image and video
Wilson Yan, Matei Zaharia, Volodymyr Mnih, Pieter Abbeel, Aleksandra Faust, and Hao Liu · 2024
Closest in time.
An image is worth 32 tokens for reconstruction and generation
Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen · 2024
Closest in time.
Scaling the codebook size of vqgan to 100,000 with a utilization rate of 99
Lei Zhu, Fangyun Wei, Yanye Lu, and Dong Chen · 2024
Closest in time.