Fetching the paper…
Reading the bibliography…
We introduce a novel sequential modeling approach which enables learning a Large Vision Model (LVM) without making use of any linguistic data.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Standardization of progressive matrices, 1938
John C Raven · 1941
Earlier work this paper cites.
A mathematical theory of communication
Claude E. Shannon · 1948
Earlier work this paper cites.
Prediction and entropy of printed english
Claude E Shannon · 1951
Earlier work this paper cites.
Some informational aspects of visual perception
Fred Attneave · 1954
Earlier work this paper cites.
Computational Models for Texture Analysis and Texture Synthesis
David Donovan Garber · 1981
Earlier work this paper cites.
Novel cluster-based probability model for texture synthesis, classification, and compression
Kris Popat and Rosalind W. Picard · 1993
Earlier work this paper cites.
Texture synthesis by non-parametric sampling
Alexei A. Efros and Thomas K. Leung · 1999
Earlier work this paper cites.
Video textures
Arno Schodl, Richard Szeliski, David H. Salesin, and Irfan Essa · 2000
Earlier work this paper cites.
Image quilting for texture synthesis and transfer
Alexei A. Efros and William T. Freeman · 2001
Earlier work this paper cites.
Image analogies
Aaron Hertzmann, Charles E Jacobs, Nuria Oliver, Brian Curless, and David H Salesin · 2001
Earlier work this paper cites.
Interactive motion generation from examples
Okan Arikan and D. A. Forsyth · 2002
Earlier work this paper cites.
Motion Graphs
Lucas Kovar, Michael Gleicher, and Frederic Pighin · 2002
Earlier work this paper cites.
Interactive control of avatars animated with human motion data
Jehee Lee, Jinxiang Chai, Paul Reitsma, Jessica K. Hodgins, and Nancy Pollard · 2002
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning to recognize objects in egocentric activities
Alireza Fathi, Xiaofeng Ren, and James M Rehg · 2011
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Ava: A large-scale database for aesthetic visual analysis
Naila Murray, Luca Marchesotti, and Florent Perronnin · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
Pradipto Das, Chenliang Xu, Richard F Doell, and Jason J Corso · 2013
Earlier work this paper cites.
Towards understanding action recognition
H. Jhuang, J. Gall, S. Zuffi, C. Schmid, and M. J. Black · 2013
Earlier work this paper cites.
2d human pose estimation: New benchmark and state of the art analysis
Mykhaylo Andriluka, Leonid Pishchulin, Peter Gehler, and Bernt Schiele · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Beyond pascal: A benchmark for 3d object detection in the wild
Yu Xiang, Roozbeh Mottaghi, and Silvio Savarese · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A. Efros · 2015
Earlier work this paper cites.
Region-based convolutional networks for accurate object detection and segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Cited alongside, same era.
Stacked hourglass networks for human pose estimation
Alejandro Newell, Kaiyu Yang, and Jia Deng · 2016
Cited alongside, same era.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Cited alongside, same era.
Context encoders: Feature learning by inpainting
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros · 2016
Cited alongside, same era.
A benchmark dataset and evaluation methodology for video object segmentation
Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung · 2016
Cited alongside, same era.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Semantic understanding of scenes through the ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2019
Later among the works it cites.
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever · 2020
Later among the works it cites.
Exploring simple siamese representation learning
Xinlei Chen and Kaiming He · 2020
Later among the works it cites.
Openmmlab pose estimation toolbox and benchmark
MMPose Contributors · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Cited alongside, same era.
Colorful image colorization
Richard Zhang, Phillip Isola, and Alexei A Efros · 2016
Cited alongside, same era.
Multi-task self-supervised visual learning
Carl Doersch and Andrew Zisserman · 2017
Cited alongside, same era.
The” something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Cited alongside, same era.
Ubernet: Training a universal convolutional neural network for low-, mid-, and high-level vision using diverse datasets and limited memory
Iasonas Kokkinos · 2017
Cited alongside, same era.
Unite the people: Closing the loop between 3d and 2d human representations
Christoph Lassner, Javier Romero, Martin Kiefel, Federica Bogo, Michael J Black, and Peter V Gehler · 2017
Cited alongside, same era.
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al · 2020
Later among the works it cites.
Dense extreme inception network: Towards a robust cnn model for edge detection
X. Soria, E. Riba, and A. Sappa · 2020
Later among the works it cites.
Estimating and exploiting the aleatoric uncertainty in surface normal estimation
Gwangbin Bae, Ignas Budvytis, and Roberto Cipolla · 2021
Later among the works it cites.
Beit: Bert pre-training of image transformers
Hangbo Bao, Li Dong, and Furu Wei · 2021
Later among the works it cites.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2021
Later among the works it cites.
Perceiver io: A general architecture for structured inputs & outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et al · 2021
Later among the works it cites.
Multisports: A multi-person video dataset of spatio-temporally localized sports actions
Yixuan Li, Lei Chen, Runyu He, Zhenzhi Wang, Gangshan Wu, and Limin Wang · 2021
Later among the works it cites.
Multi-moments in time: Learning and interpreting models for multi-action video understanding
Mathew Monfort, Bowen Pan, Kandan Ramakrishnan, Alex Andonian, Barry A McNamara, Alex Lascelles, Quanfu Fan, Dan Gutfreund, Rogério Schmidt Feris, and Aude Oliva · 2021
Later among the works it cites.
Vision transformers for dense prediction
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun · 2021
Later among the works it cites.
Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction
Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny · 2021
Later among the works it cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki · 2021
Later among the works it cites.
Simmim: A simple framework for masked image modeling
Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu · 2021
Later among the works it cites.
Visual prompting via image inpainting
Amir Bar, Yossi Gandelsman, Trevor Darrell, Amir Globerson, and Alexei Efros · 2022
Later among the works it cites.
Maskgit: Masked generative image transformer
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman · 2022
Later among the works it cites.
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Later among the works it cites.
Large-scale video panoptic segmentation in the wild: A benchmark
Jiaxu Miao, Xiaohan Wang, Yu Wu, Wei Li, Xu Zhang, Yunchao Wei, and Yi Yang · 2022
Later among the works it cites.
Scaling autoregressive models for content-rich text-to-image generation, 2022
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu · 2022
Later among the works it cites.
Instructpix2pix: Learning to follow image editing instructions
Tim Brooks, Aleksander Holynski, and Alexei A. Efros · 2023
Closest in time.
Objaverse: A universe of annotated 3d objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Images speak in images: A generalist painter for in-context visual learning
Xinlong Wang, Wen Wang, Yue Cao, Chunhua Shen, and Tiejun Huang · 2023
Closest in time.
On the de-duplication of laion-2b, 2023
Ryan Webster, Julien Rabin, Loic Simon, and Frederic Jurie · 2023
Closest in time.