Fetching the paper…
Reading the bibliography…
Efficient video tokenization remains a key bottleneck in learning general purpose vision models that are capable of processing long video sequences.
The jpeg2000 still image coding system: an overview
Charilaos Christopoulos, Athanassios Skodras, and Touradj Ebrahimi · 2000
Earlier work this paper cites.
Overview of the h. 264/avc video coding standard
Thomas Wiegand, Gary J Sullivan, Gisle Bjontegaard, and Ajay Luthra · 2003
Earlier work this paper cites.
H. 264 and MPEG-4 video compression: video coding for next-generation multimedia
Iain E Richardson · 2004
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Learning ordered representations with nested dropout
Oren Rippel, Michael Gelbart, and Ryan Adams · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
High efficiency video coding (hevc)
Vivienne Sze, Madhukar Budagavi, and Gary J Sullivan · 2014
Earlier work this paper cites.
A technical overview of vp9—the latest open-source video codec
Debargha Mukherjee, Jingning Han, Jim Bankoski, Ronald Bultje, Adrian Grange, John Koleszar, Paul Wilkins, and Yaowu Xu · 2015
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei · 2016
Earlier work this paper cites.
Deepcoder: A deep neural network based video compression
Tong Chen, Haojie Liu, Qiu Shen, Tao Yue, Xun Cao, and Zhan Ma · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Anytime neural prediction via slicing networks vertically
Hankook Lee and Jinwoo Shin · 2018
Earlier work this paper cites.
Joint autoregressive and hierarchical priors for learned image compression
David Minnen, Johannes Ballé, and George D Toderici · 2018
Earlier work this paper cites.
Jiahui Yu, Linjie Yang, Ning Xu, Jianchao Yang, and Thomas Huang · 2018
Cited alongside, same era.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin · 2019
Cited alongside, same era.
Generating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron Van den Oord, and Oriol Vinyals · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy · 2020
Cited alongside, same era.
Dynabert: Dynamic bert with adaptive width and depth
Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu · 2020
Cited alongside, same era.
Transframer: Arbitrary frame prediction with generative models
Charlie Nash, Joao Carreira, Jacob Walker, Iain Barr, Andrew Jaegle, Mateusz Malinowski, and Peter Battaglia · 2022
Later among the works it cites.
Lossy image compression with compressive autoencoders
Lucas Theis, Wenzhe Shi, Andrew Cunningham, and Ferenc Huszár · 2022
Later among the works it cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai · 2023
Later among the works it cites.
Sharegpt4v: Improving large multi-modal models with better captions
Lin Chen, Jisong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin · 2023
Later among the works it cites.
Openllama: An open reproduction of llama
Xinyang Geng and Hao Liu · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gyuwan Kim and Kyunghyun Cho · 2020
Cited alongside, same era.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Cited alongside, same era.
Variable-rate discrete representation learning
Sander Dieleman, Charlie Nash, Jesse Engel, and Karen Simonyan · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Cited alongside, same era.
A technical overview of av1
Jingning Han, Bohan Li, Debargha Mukherjee, Ching-Han Chiang, Adrian Grange, Cheng Chen, Hui Su, Sarah Parker, Sai Deng, Urvang Joshi, et al · 2021
Cited alongside, same era.
Deep contextual video compression
Jiahao Li, Bin Li, and Yan Lu · 2021
Cited alongside, same era.
Videogpt: Video generation using vq-vae and transformers
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas · 2021
Cited alongside, same era.
Later among the works it cites.
Video-llava: Learning united visual representation by alignment before projection
Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan · 2023
Later among the works it cites.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Later among the works it cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan · 2023
Later among the works it cites.
Finite scalar quantization: Vq-vae made simple
Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Later among the works it cites.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh · 2024
Closest in time.
Finite scalar quantization: VQ-VAE made simple
Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen · 2024
Closest in time.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang · 2024
Closest in time.
Emu3: Next-token prediction is all you need
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, et al · 2024
Closest in time.
Show-o: One single transformer to unify multimodal understanding and generation
Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, Zhenheng Yang, and Mike Zheng Shou · 2024
Closest in time.