Fetching the paper…
Reading the bibliography…
We show how transformers can be used to vastly simplify neural video compression.
“Language models are few-shot learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“Multiscale structural similarity for image quality assessment”
Zhou Wang, Eero Simoncelli and Alan Bovik · 2003
Earlier work this paper cites.
“Overview of the high efficiency video coding (HEVC) standard”
Gary Sullivan, Jens-Rainer Ohm, Woo-Jin Han and Thomas Wiegand · 2012
Earlier work this paper cites.
“Auto-encoding variational bayes”
Diederik Kingma and Max Welling · 2013
Earlier work this paper cites.
“MCL-JCV: a JND-based H.264/AVC video quality assessment dataset”
Haiqiang Wang et al · 2016
Earlier work this paper cites.
“Lossy image compression with compressive autoencoders”
Lucas Theis, Wenzhe Shi, Andrew Cunningham and Ferenc Huszar · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Variational image compression with a scale hyperprior”
Johannes Ballé et al · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Earlier work this paper cites.
“Understanding back-translation at scale”
Sergey Edunov, Myle Ott, Michael Auli and David Grangier · 2018
Earlier work this paper cites.
“Joint autoregressive and hierarchical priors for learned image compression”
David Minnen, Johannes Ballé and George Toderici · 2018
Earlier work this paper cites.
“Video compression through image interpolation”
Chao-Yuan Wu, Nayan Singhal and Philipp Krahenbuhl · 2018
Earlier work this paper cites.
“Neural inter-frame compression for video coding”
Abdelaziz Djelouah, Joaquim Campos, Simone Schaub-Meyer and Christopher Schroers · 2019
Earlier work this paper cites.
“Video compression with rate-distortion autoencoders”
Amirhossein Habibian, Ties Rozendaal, Jakub Tomczak and Taco Cohen · 2019
Earlier work this paper cites.
“Dvc: An end-to-end deep video compression framework”
Guo Lu et al · 2019
Earlier work this paper cites.
“Videobert: A joint model for video and language representation learning”
Chen Sun et al · 2019
Earlier work this paper cites.
“Making convolutional networks shift-invariant again”
Richard Zhang · 2019
Cited alongside, same era.
“Scale-space flow for end-to-end optimized video compression”
Eirikur Agustsson et al · 2020
Cited alongside, same era.
“Nonlinear transform coding”
Johannes Ballé et al · 2020
Cited alongside, same era.
“An image is worth 16x16 words: Transformers for image recognition at scale”
Alexey Dosovitskiy et al · 2020
Cited alongside, same era.
“Feedback recurrent autoencoder for video compression”
Adam Golinski et al · 2020
Cited alongside, same era.
“Flax: A neural network library and ecosystem for JAX”, 2020
Jonathan Heek et al · 2020
Cited alongside, same era.
“Deep contextual video compression”
Jiahao Li, Bin Li and Yan Lu · 2021
Later among the works it cites.
“Deep learning in latent space for video prediction and compression”
Bowen Liu, Yu Chen, Shiyu Liu and Hun-Seok Kim · 2021
Later among the works it cites.
“Swin transformer: Hierarchical vision transformer using shifted windows”
Ze Liu et al · 2021
Later among the works it cites.
“Neural Video Compression using GANs for Detail Synthesis and Propagation”
Fabian Mentzer et al · 2021
Later among the works it cites.
“Video transformer network”
Daniel Neimark, Omri Bar, Maya Zohar and Dotan Asselmann · 2021
Later among the works it cites.
“Elf-vc: Efficient learned flexible-rate video coding”
Oren Rippel et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Conditional entropy coding for efficient video compression”
Jerry Liu et al · 2020
Cited alongside, same era.
“UVG dataset: 50/120fps 4K sequences for video codec analysis and development”
Alexandre Mercat, Marko Viitanen and Jarno Vanne · 2020
Cited alongside, same era.
“Channel-wise autoregressive entropy models for learned image compression”
David Minnen and Saurabh Singh · 2020
Cited alongside, same era.
“CLIC 2020: Challenge on Learned Image Compression”
George Toderici et al · 2020
Cited alongside, same era.
“Hierarchical autoregressive modeling for neural video compression”
Ruihan Yang, Yibo Yang, Joseph Marino and Stephan Mandt · 2020
Cited alongside, same era.
“Vivit: A video vision transformer”
Anurag Arnab et al · 2021
Cited alongside, same era.
“Max-deeplab: End-to-end panoptic segmentation with mask transformers”
Huiyu Wang et al · 2021
Later among the works it cites.
“Perceptual Learned Video Compression with Recurrent Conditional GAN”
Ren Yang, Luc Van and Radu Timofte · 2021
Later among the works it cites.
“Learning for Video Compression with Recurrent Auto-Encoder and Recurrent Probability Model”
Ren Yang, Fabian Mentzer, Luc Van and Radu Timofte · 2021
Later among the works it cites.
“Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers”
Sixiao Zheng et al · 2021
Later among the works it cites.
“Transformer-based Transform Coding”
Yinhao Zhu, Yang Yang and Taco Cohen · 2021
Later among the works it cites.
“Palm: Scaling language modeling with pathways”
Aakanksha Chowdhery et al · 2022
Closest in time.
“Elic: Efficient learned image compression with unevenly grouped space-channel contextual adaptive coding”
Dailan He et al · 2022
Closest in time.
“Entroformer: A Transformer-based Entropy Model for Learned Image Compression”
Yichen Qian et al · 2022
Closest in time.
“An Introduction to Neural Data Compression” preprint, 2022
Y. Yang, S. Mandt and L. Theis · 2022
Closest in time.