Fetching the paper…
Reading the bibliography…
In this work, we reveal the limitations of visual tokenizers and VAEs in preserving fine-grained features, and propose a benchmark to evaluate reconstruction performance for two challenging visual contents: text and face.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli · 2004
Earlier work this paper cites.
Labeled faces in the wild: A database forstudying face recognition in unconstrained environments
Gary B Huang, Marwan Mattar, Tamara Berg, and Eric Learned-Miller · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization
Martin Koestinger, Paul Wohlhart, Peter M Roth, and Horst Bischof · 2011
Earlier work this paper cites.
Hdr-vdp-2: A calibrated visual metric for visibility and quality predictions in all luminance conditions
Rafat Mantiuk, Kil Joong Kim, Allan G Rempel, and Wolfgang Heidrich · 2011
Earlier work this paper cites.
Fsim: A feature similarity index for image quality assessment
Lin Zhang, Lei Zhang, Xuanqin Mou, and David Zhang · 2011
Earlier work this paper cites.
Icdar 2013 robust reading competition
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere De Las Heras · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Icdar 2015 competition on robust reading
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, et al · 2015
Earlier work this paper cites.
300 faces in-the-wild challenge: Database and results
Christos Sagonas, Epameinondas Antonakos, Georgios Tzimiropoulos, Stefanos Zafeiriou, and Maja Pantic · 2016
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Frontal to profile face verification in the wild
Soumyadip Sengupta, Jun-Cheng Chen, Carlos Castillo, Vishal M. Patel, Rama Chellappa, and David W. Jacobs · 2016
Earlier work this paper cites.
Coco-text: Dataset and benchmark for text detection and recognition in natural images
Andreas Veit, Tomas Matera, Lukas Neumann, Jiri Matas, and Serge Belongie · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Icdar2017 robust reading challenge on multi-lingual scene text detection and script identification-rrc-mlt
Nibal Nayef, Fei Yin, Imen Bizid, Hyunsoo Choi, Yuan Feng, Dimosthenis Karatzas, Zhenbo Luo, Umapada Pal, Christophe Rigaud, Joseph Chazalon, et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Cross-age LFW: A database for studying cross-age face recognition in unconstrained environments
Tianyue Zheng, Weihong Deng, and Jiani Hu · 2017
Earlier work this paper cites.
Look at boundary: A boundary-aware face alignment algorithm
Wayne Wu, Chen Qian, Shuo Yang, Quan Wang, Yici Cai, and Qiang Zhou · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
Diagnosing and enhancing vae models
Bin Dai and David Wipf · 2019
Earlier work this paper cites.
Icdar2019 competition on scanned receipt ocr and information extraction
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and CV Jawahar · 2019
Earlier work this paper cites.
Curved scene text detection via transverse and longitudinal sequence connection
Yuliang Liu, Lianwen Jin, Shuaitao Zhang, Canjie Luo, and Sheng Zhang · 2019
Earlier work this paper cites.
Cord: a consolidated receipt dataset for post-ocr parsing
Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee · 2019
Earlier work this paper cites.
Icdar 2019 robust reading challenge on reading chinese text on signboard
Rui Zhang, Yongsheng Zhou, Qianyi Jiang, Qi Song, Nan Li, Kai Zhou, Lei Wang, Dong Wang, Minghui Liao, Mingkun Yang, et al · 2019
Earlier work this paper cites.
Total-text: toward orientation robustness in scene text detection
Chee-Kheng Chng, Chee Seng Chan, and Cheng-Lin Liu · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Abcnet: Real-time scene text spotting with adaptive bezier-curve network
Yuliang Liu, Hao Chen, Chunhua Shen, Tong He, Lianwen Jin, and Liangwei Wang · 2020
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Cited alongside, same era.
Docvqa: A dataset for vqa on document images
Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar · 2021
Cited alongside, same era.
doctr: Document text recognition
Mindee · 2021
Cited alongside, same era.
Open-sora plan: Open-source large video generation model
Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Liuhan Chen, et al · 2024
Later among the works it cites.
World model on million-length video and language with ringattention
Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel · 2024
Later among the works it cites.
Open-magvit2: An open-source project toward democratizing auto-regressive visual generation
Zhuoyan Luo, Fengyuan Shi, Yixiao Ge, Yujiu Yang, Limin Wang, and Ying Shan · 2024
Later among the works it cites.
Cosmos-tokenizer
NVIDIA · 2024
Later among the works it cites.
Tokenflow: Unified image tokenizer for multimodal understanding and generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Cited alongside, same era.
Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text
Amanpreet Singh, Guan Pang, Mandy Toh, Jing Huang, Wojciech Galuba, and Tal Hassner · 2021
Cited alongside, same era.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2021
Cited alongside, same era.
Scene text recognition with permuted autoregressive sequence models
Darwin Bautista and Rowel Atienza · 2022
Cited alongside, same era.
Infographicvqa
Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Liao Qu, Huichao Zhang, Yiheng Liu, Xu Wang, Yi Jiang, Yiming Gao, Hu Ye, Daniel K Du, Zehuan Yuan, and Xinglong Wu · 2024
Later among the works it cites.
Autoregressive model beats diffusion: Llama for scalable image generation
Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan · 2024
Later among the works it cites.
Chameleon: Mixed-modal early-fusion foundation models
Chameleon Team · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang · 2024
Later among the works it cites.
Omnitokenizer: A joint image-video tokenizer for visual generation
Junke Wang, Yi Jiang, Zehuan Yuan, Binyue Peng, Zuxuan Wu, and Yu-Gang Jiang · 2024
Later among the works it cites.
Maskbit: Embedding-free image generation via bit tokens
Mark Weber, Lijun Yu, Qihang Yu, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen · 2024
Later among the works it cites.
Liquid: Language models are scalable multi-modal generators
Junfeng Wu, Yi Jiang, Chuofan Ma, Yuliang Liu, Hengshuang Zhao, Zehuan Yuan, Song Bai, and Xiang Bai · 2024
Later among the works it cites.
Dstext v2: A comprehensive video text spotting dataset for dense and small text
Weijia Wu, Yiming Zhang, Yefei He, Luoming Zhang, Zhenyu Lou, Hong Zhou, and Xiang Bai · 2024
Later among the works it cites.
Sana: Efficient high-resolution image synthesis with linear diffusion transformers
Enze Xie, Junsong Chen, Junyu Chen, Han Cai, Haotian Tang, Yujun Lin, Zhekai Zhang, Muyang Li, Ligeng Zhu, Yao Lu, et al · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al · 2024
Later among the works it cites.
An image is worth 32 tokens for reconstruction and generation
Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen · 2024
Later among the works it cites.
Image and video tokenization with binary spherical quantization
Yue Zhao, Yuanjun Xiong, and Philipp Krähenbühl · 2024
Later among the works it cites.
Open-sora: Democratizing efficient video production for all
Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You · 2024
Later among the works it cites.
FlexTok: Resampling images into 1d token sequences of flexible length
Roman Bachmann, Jesse Allardice, David Mizrahi, Enrico Fini, Oğuzhan Fatih Kar, Elmira Amirloo, Alaaeldin El-Nouby, Amir Zamir, and Afshin Dehghan · 2025
Closest in time.
Unitok: A unified tokenizer for visual generation and understanding
Chuofan Ma, Yi Jiang, Junfeng Wu, Jihan Yang, Xin Yu, Zehuan Yuan, Bingyue Peng, and Xiaojuan Qi · 2025
Closest in time.
Step-video-t2v technical report: The practice, challenges, and future of video foundation model
Guoqing Ma, Haoyang Huang, Kun Yan, Liangyu Chen, Nan Duan, Shengming Yin, Changyi Wan, Ranchen Ming, Xiaoniu Song, Xing Chen, et al · 2025
Closest in time.
Bridging continuous and discrete tokens for autoregressive visual generation
Yuqing Wang, Zhijie Lin, Yao Teng, Yuanzhi Zhu, Shuhuai Ren, Jiashi Feng, and Xihui Liu · 2025
Closest in time.
Qingsong Xie, Zhao Zhang, Zhe Huang, Yanhao Zhang, Haonan Lu, and Zhenyu Yang · 2025
Closest in time.
Reconstruction vs. generation: Taming optimization dilemma in latent diffusion models
Jingfeng Yao, Bin Yang, and Xinggang Wang · 2025
Closest in time.
Representation alignment for generation: Training diffusion transformers is easier than you think
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie · 2025
Closest in time.