Fetching the paper…
Reading the bibliography…
With the introduction of transformer-based models for vision and language tasks, such as LLaVA and Chameleon, there has been renewed interest in the discrete tokenized representation of images.
Relative frequency of English speech sounds
Godfrey Dewey · 1921
Earlier work this paper cites.
Some statistics of evolution and geographical distribution in plants and animals, and their significance
John C Willis and G Udny Yule · 1922
Earlier work this paper cites.
Selected studies of the principle of relative frequency in language
George Kingsley Zipf · 1932
Earlier work this paper cites.
The law of anomalous numbers
Frank Benford · 1938
Earlier work this paper cites.
Prediction and entropy of printed english
Claude E Shannon · 1951
Earlier work this paper cites.
A method for the construction of minimum-redundancy codes
David A Huffman · 1952
Earlier work this paper cites.
Contribution à la théorie mathématique des jeux de communication
Benoît Mandelbrot · 1953
Earlier work this paper cites.
On a class of skew distribution functions
Herbert A Simon · 1955
Earlier work this paper cites.
An Empirical Bayes Approach to Statistics
Herbert Robbins · 1956
Earlier work this paper cites.
The algebraic theory of context-free languages
Noam Chomsky and Marcel P Schützenberger · 1959
Earlier work this paper cites.
Quantitative linguistics
Gustav Herdan · 1964
Earlier work this paper cites.
Generalized procrustes analysis
John C Gower · 1975
Earlier work this paper cites.
Information retrieval: Computational and theoretical aspects
Harold Stanley Heaps · 1978
Earlier work this paper cites.
Trainable grammars for speech recognition
James K Baker · 1979
Earlier work this paper cites.
Hausdorff dimension of quasi-circles
Rufus Bowen · 1979
Earlier work this paper cites.
Origins of scaling in natural images
Daniel L Ruderman · 1997
Earlier work this paper cites.
Quantitative tools for comparing animal communication systems: information theory applied to bottlenose dolphin whistle repertoires
Brenda McCowan, Sean F Hanser, and Laurance R Doyle · 1999
Earlier work this paper cites.
Zipf and heaps laws’ coefficients depend on language
Alexander Gelbukh and Grigori Sidorov · 2001
Earlier work this paper cites.
Zipf’s law in image coding schemes
Michael Crosier and Lewis D Griffin · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li · 2009
Earlier work this paper cites.
Benford’s law as an indicator of fraud in economics
Karl-Heinz Tödter · 2009
Earlier work this paper cites.
Benford’s law in linguistic texts: Its principle and applications
Jung-Ha Hong · 2010
Earlier work this paper cites.
Benford’s law in the natural sciences
Malcolm Sambridge, Hrvoje Tkalčić, and A Jackson · 2010
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma · 2013
Cited alongside, same era.
powerlaw: a python package for analysis of heavy-tailed distributions
Jeff Alstott, Ed Bullmore, and Dietmar Plenz · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Cited alongside, same era.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Cited alongside, same era.
Benford’s law
Steven J Miller · 2015
Cited alongside, same era.
CLIPScore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Later among the works it cites.
Hubert: How much can a bad teacher benefit asr pre-training?
Wei-Ning Hsu, Yao-Hung Hubert Tsai, Benjamin Bolte, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Later among the works it cites.
Discrete representations strengthen vision transformer robustness
Chengzhi Mao, Lu Jiang, Mostafa Dehghani, Carl Vondrick, Rahul Sukthankar, and Irfan Essa · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
What’s in a caption? dataset-specific linguistic diversity and its effect on visual description models and metrics
David M. Chan, Austin Myers, Sudheendra Vijayanarasimhan, David A. Ross, Bryan Seybold, and John F. Canny · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Cited alongside, same era.
Word2vec
Kenneth Ward Church · 2017
Cited alongside, same era.
Zipf’s and benford’s laws in twitter hashtags
José Alberto Pérez Melián, J Alberto Conejero, and Cesar Ferri Ramirez · 2017
Cited alongside, same era.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Cited alongside, same era.
Constituency parsing with a self-attentive encoder
Nikita Kitaev and Dan Klein · 2018
Cited alongside, same era.
Later among the works it cites.
Make-a-scene: Scene-based text-to-image generation with human priors
Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman · 2022
Later among the works it cites.
Vector quantized diffusion model for text-to-image synthesis
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Crossmodal-3600: A massively multilingual multimodal evaluation dataset
Ashish V Thapliyal, Jordi Pont Tuset, Xi Chen, and Radu Soricut · 2022
Later among the works it cites.
Benford’s law applies to word frequency rank in english, german, french, spanish, and italian
Jennifer Golbeck · 2023
Later among the works it cites.
Unsupervised discontinuous constituency parsing with mildly context-sensitive grammars
Songlin Yang, Roger Levy, and Yoon Kim · 2023
Later among the works it cites.
Sequential modeling enables scalable learning for large vision models
Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan L Yuille, Trevor Darrell, Jitendra Malik, and Alexei A Efros · 2024
Closest in time.
Re-evaluating the need for visual signals in unsupervised grammar induction
Boyi Li, Rodolfo Corona, Karttikeya Mangalam, Catherine Chen, Daniel Flaherty, Serge Belongie, Kilian Weinberger, Jitendra Malik, Trevor Darrell, and Dan Klein · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Spin: Hierarchical segmentation with subpart granularity in natural images
Josh Myers-Dean, Jarek Reynolds, Brian Price, Yifei Fan, and Danna Gurari · 2024
Closest in time.
Autoregressive model beats diffusion: Llama for scalable image generation
Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan · 2024
Closest in time.
Chameleon: Mixed-modal early-fusion foundation models
Chameleon Team · 2024
Closest in time.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et al · 2024
Closest in time.
mplug-owl3: Towards long image-sequence understanding in multi-modal large language models
Jiabo Ye, Haiyang Xu, Haowei Liu, Anwen Hu, Ming Yan, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou · 2024
Closest in time.
Transfusion: Predict the next token and diffuse images with one multi-modal model
Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, and Omer Levy · 2024
Closest in time.