An image is worth 16x16 words: Transformers for image recognition at scale
Original
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Flax: A neural network library and ecosystem for JAX, 2020
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee · 2020
Cited alongside, same era.
Exploring the limits of large scale pre-training
Original
Samira Abnar, Mostafa Dehghani, Behnam Neyshabur, and Hanie Sedghi · 2021
Cited alongside, same era.
ViViT: A video vision transformer
Original
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Cited alongside, same era.
Attention bottlenecks for multimodal fusion
Original
Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun · 2021
Cited alongside, same era.
http://github.com/google/CommonLoopUtils
CLU - common loop utils
Cited in the paper.
https://github.com/deepmind/dmvr
DMVR: Deepmind video readers
Cited in the paper.
https://github.com/google-research/ott
Optimal Transport Tools (OTT), a toolbox for everything wasserstein
Cited in the paper.
https://www.tensorflow.org/datasets
TensorFlow Datasets, a collection of ready-to-use datasets
Cited in the paper.