Fetching the paper…
Reading the bibliography…
Recent work in visual representation learning for robotics demonstrates the viability of learning from large video datasets of humans performing everyday tasks.
Dynamic sensor-based control of robots with visual feedback
Lee E. Weiss, Arthur C. Sanderson, and Charles P. Neuman. 1987 · 1987
Earlier work this paper cites.
Visual servo control. I. Basic approaches
François Chaumette and Seth A. Hutchinson. 2006 · 2006
Earlier work this paper cites.
Cost-Based Anticipatory Action Selection for Human–Robot Fluency
Guy Hoffman and Cynthia Breazeal. 2007 · 2007
Earlier work this paper cites.
Robotic Grasping of Novel Objects using Vision
Ashutosh Saxena, Justin Driemeyer, and A. Ng. 2008 · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database. In Computer Vision and Pattern Recognition (CVPR) . 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Understanding Natural Language Commands for Robotic Navigation and Mobile Manipulation. In Association for the Advancement of Artificial Intelligence (AAAI)
Stefanie Tellex, Thomas Kollar, Steven Dickerson, Matthew R Walter, Ashis Gopal Banerjee, Seth J Teller, and Nicholas Roy. 2011 · 2011
Earlier work this paper cites.
Are we ready for autonomous driving? The KITTI vision benchmark suite. In Computer Vision and Pattern Recognition (CVPR) . 3354–3361
Andreas Geiger, Philip Lenz, and Raquel Urtasun. 2012 · 2012
Earlier work this paper cites.
Recognition, prediction, and planning for assisted teleoperation of freeform tasks
Kris K. Hauser. 2012 · 2012
Earlier work this paper cites.
Intention-Aware Motion Planning. In Workshop for the Algorithmic Foundations of Robotics (WAFR)
Tirthankar Bandyopadhyay, Kok Sung Won, Emilio Frazzoli, David Hsu, Wee Sun Lee, and Daniela Rus. 2013 · 2013
Earlier work this paper cites.
Data-Driven Grasp Synthesis—A Survey
Jeannette Bohg, Antonio Morales, Tamim Asfour, and Danica Kragic. 2013 · 2013
Earlier work this paper cites.
Decoding with Large-Scale Neural Language Models Improves Translation. In Empirical Methods in Natural Language Processing (EMNLP) . 1387–1392
Ashish Vaswani, Yinggong Zhao, Victoria Fossum, and David Chiang. 2013 · 2013
Earlier work this paper cites.
Microsoft COCO: Common objects in context. In European Conference on Computer Vision (ECCV) . 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In International Conference on Machine Learning (ICML) . 448–456
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Learning state representations with robotic priors
Rico Jonschkowski and Oliver Brock. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR)
Diederik Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop
Fisher Yu, Yinda Zhang, Shuran Song, Ari Seff, and Jianxiong Xiao. 2015 · 2015
Earlier work this paper cites.
Analysis and Observations From the First Amazon Picking Challenge
Nikolaus Correll, Kostas E. Bekris, Dmitry Berenson, Oliver Brock, Albert J. Causo, Kris K. Hauser, Kei Okada, Alberto Rodriguez, Joseph M. Romano, and Peter R. Wurman. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In Computer Vision and Pattern Recognition (CVPR)
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Gaussian Error Linear Units (GELUs)
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
End-to-End Training of Deep Visuomotor Policies
S. Levine, Chelsea Finn, Trevor Darrell, and P. Abbeel. 2016 · 2016
Earlier work this paper cites.
Modeling Context in Referring Expressions. In European Conference on Computer Vision (ECCV)
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg. 2016 · 2016
Earlier work this paper cites.
Accurately and Efficiently Interpreting Human-Robot Instructions of Varying Granularities. In Robotics: Science and Systems (RSS)
Dilip Arumugam, Siddharth Karamcheti, Nakul Gopalan, Lawson L. S. Wong, and Stefanie Tellex. 2017 · 2017
Earlier work this paper cites.
The “Something Something Video Database for Learning and Evaluating Visual Common Sense. In International Conference on Computer Vision (ICCV)
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fründ, Peter N. Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, and Roland Memisevic. 2017 · 2017
Earlier work this paper cites.
Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics. In Robotics: Science and Systems (RSS)
Jeffrey Mahler, Jacky Liang, Sherdil Niyaz, Michael Laskey, Richard Doan, Xinyu Liu, Juan Aparicio Ojea, and Ken Goldberg. 2017 · 2017
Earlier work this paper cites.
Mapping Instructions and Visual Observations to Actions with Reinforcement Learning. In Empirical Methods in Natural Language Processing (EMNLP)
Dipendra K. Misra, John Langford, and Yoav Artzi. 2017 · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Robotic pick-and-place of novel objects in clutter with multi-affordance grasping and cross-domain image matching
Andy Zeng, Shuran Song, Kuan-Ting Yu, Elliott Donlon, Francois Robert Hogan, Maria Bauzá, Daolin Ma, Orion Taylor, Melody Liu, Eudald Romo, Nima Fazeli, Ferran Alet, Nikhil Chavan Dafle, Rachel Holladay, Isabella Morona, Prem Qu Nair, Druck Green, Ian Taylor, Weber Liu, Thomas A. Funkhouser, and Alberto Rodriguez. 2017 · 2017
Earlier work this paper cites.
Scaling Egocentric Vision: The EPIC-KITCHENS Dataset. In European Conference on Computer Vision (ECCV)
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. 2018 · 2018
Earlier work this paper cites.
Shared autonomy via hindsight optimization for teleoperation and teaming
Shervin Javdani, Henny Admoni, Stefania Pellegrinelli, Siddhartha S Srinivasa, and J Andrew Bagnell. 2018 · 2018
Earlier work this paper cites.
Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks. In International Conference on Machine Learning (ICML)
Juho Lee, Yoonho Lee, Jungtaek Kim, Adam R. Kosiorek, Seungjin Choi, and Yee Whye Teh. 2018 · 2018
Earlier work this paper cites.
Time-Contrastive Networks: Self-Supervised Learning from Video. In International Conference on Robotics and Automation (ICRA) . 1134–1141
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, and Sergey Levine. 2018 · 2018
Earlier work this paper cites.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In Association for Computational Linguistics (ACL)
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Association for Computational Linguistics (ACL) . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
DeepMDP: Learning Continuous Latent Space Models for Representation Learning. In International Conference on Machine Learning (ICML)
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G. Bellemare. 2019 · 2019
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. In Advances in Neural Information Processing Systems (NeurIPS)
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 2019
Cited alongside, same era.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Cited alongside, same era.
PyTorch Image Models
Ross Wightman. 2019 · 2019
Cited alongside, same era.
Root Mean Square Layer Normalization. In Advances in Neural Information Processing Systems (NeurIPS)
Biao Zhang and Rico Sennrich. 2019 · 2019
OCID-Ref: A 3D Robotic Dataset With Embodied Language For Clutter Scene Grounding. In Association for Computational Linguistics (ACL)
Ke-Jyun Wang, Yun-Hsuan Liu, Hung-Ting Su, Jen-Wei Wang, Yu-Siang Wang, Winston H. Hsu, and Wen-Chin Chen. 2021 · 2021
Later among the works it cites.
Learning Invariant Representations for Reinforcement Learning without Reconstruction. In International Conference on Learning Representations (ICLR)
Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine. 2021 · 2021
Later among the works it cites.
Rethinking Semantic Segmentation from a Sequence-to-Sequence perspective with Transformers. In Computer Vision and Pattern Recognition (CVPR)
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip H. S. Torr, and Li Zhang. 2021 · 2021
Later among the works it cites.
CM3: A Causal Masked Multimodal Model of the Internet
Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, and Luke Zettlemoyer. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Longformer: The Long-Document Transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning (ICML) . 1597–1607
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020 · 2020
Cited alongside, same era.
Dream to Control: Learning Behaviors by Latent Imagination. In International Conference on Learning Representations (ICLR)
Danijar Hafner, Timothy P. Lillicrap, Jimmy Ba, and Mohammad Norouzi. 2020 · 2020
Cited alongside, same era.
Reinforcement Learning with Augmented Data. In Advances in Neural Information Processing Systems (NeurIPS)
Michael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, P. Abbeel, and A. Srinivas. 2020 · 2020
Cited alongside, same era.
Corey Lynch and Pierre Sermanet. 2020 · 2020
Cited alongside, same era.
Understanding Human Hands in Contact at Internet Scale. In Computer Vision and Pattern Recognition (CVPR) . 9866–9875
Dandan Shan, Jiaqi Geng, Michelle Shu, and David F. Fouhey. 2020 · 2020
Cited alongside, same era.
Concept2Robot: Learning Manipulation Concepts from Instructions and Human Demonstrations. In Robotics: Science and Systems (RSS)
Lin Shao, Toki Migimatsu, Q. Zhang, Karen Yang, and Jeannette Bohg. 2020 · 2020
Cited alongside, same era.
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alexander Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil Jayant Joshi, Ryan C. Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego M Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, and Mengyuan Yan. 2022 · 2022
Later among the works it cites.
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. 2022 · 2022
Later among the works it cites.
Human-to-Robot Imitation in the Wild. In Robotics: Science and Systems (RSS)
Shikhar Bahl, Abhi Gupta, and Deepak Pathak. 2022 · 2022
Later among the works it cites.
BEiT: BERT Pre-Training of Image Transformers. In International Conference on Learning Representations (ICLR)
Hangbo Bao, Li Dong, and Furu Wei. 2022 · 2022
Later among the works it cites.
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models. In Association for Computational Linguistics (ACL)
Elad Ben-Zaken, Shauli Ravfogel, and Yoav Goldberg. 2022 · 2022
Later among the works it cites.
PaLM: Scaling Language Modeling with Pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, A. Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, B. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, M. Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, S. Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, D. Luan, Hyeontaek Lim, Barret Zoph, A. Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, T. S. Pillai, Marie Pellat, Aitor Lewkowycz, E. Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, K. Meier-Hellstern, D. Eck, J. Dean, Slav Petrov, and Noah Fiedel. 2022 · 2022
Later among the works it cites.
Can Foundation Models Perform Zero-Shot Task Specification For Robot Manipulation?. In Learning for Dynamics & Control Conference (L4DC)
Yuchen Cui, Scott Niekum, Abhi Gupta, Vikash Kumar, and Aravind Rajeswaran. 2022 · 2022
Later among the works it cites.
Multimodal Masked Autoencoders Learn Transferable Representations
Xinyang Geng, Hao Liu, Lisa Lee, Dale Schuurams, Sergey Levine, and P. Abbeel. 2022 · 2022
Later among the works it cites.
Ego4D: Around the World in 3,000 Hours of Egocentric Video. In Computer Vision and Pattern Recognition (CVPR)
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Q. Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, F. Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Z. Xu, Chen Zhao, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Christian Fuegen, Abrham Gebreselasie, Cristina González, James M. Hillis, Xuhua Huang, Yifei Huang, Wenqi Jia, Weslie Yu Heng Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar, Federico Landini, Chao Li, Yanghao Li, Zhenqiang Li, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Yuchen Wang, Xindi Wu, Takuma Yagi, Yunyi Zhu, Pablo Arbeláez, David J. Crandall, Dima Damen, Giovanni Maria Farinella, Bernard Ghanem, Vamsi Krishna Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Kitani, Haizhou Li, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato, Jianbo Shi, Mike Zheng Shou, Antonio Torralba, Lorenzo Torresani, Mingfei Yan, and Jitendra Malik. 2022 · 2022
Later among the works it cites.
Masked Autoencoders Are Scalable Vision Learners. In Computer Vision and Pattern Recognition (CVPR)
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross B. Girshick. 2022 · 2022
Later among the works it cites.
Q-Attention: Enabling Efficient Learning for Vision-based Robotic Manipulation
Stephen James and Andrew J. Davison. 2022 · 2022
Later among the works it cites.
Coarse-to-Fine Q-Attention: Efficient Learning for Visual Robotic Manipulation via Discretisation. In Computer Vision and Pattern Recognition (CVPR) . 13729–13738
Stephen James, Kentaro Wada, Tristan Laidlow, and Andrew J. Davison. 2022 · 2022
Later among the works it cites.
InstructRL: Simple yet Effective Instruction-Following Agents with Multimodal Transformer
Hao Liu, Lisa Lee, Kimin Lee, and Pieter Abbeel. 2022 · 2022
Later among the works it cites.
VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang. 2022 · 2022
Later among the works it cites.
R3M: A Universal Visual Representation for Robot Manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta. 2022 · 2022
Later among the works it cites.
ChatGPT: Optimizing Language Models for Dialogue
OpenAI. 2022 · 2022
Later among the works it cites.
The Surprising Effectiveness of Representation Learning for Visual Imitation. In Robotics: Science and Systems (RSS)
Jyothish Pari, Nur Muhammad (Mahi) Shafiullah, Sridhar Pandian Arunachalam, and Lerrel Pinto. 2022 · 2022
Later among the works it cites.
The Unsurprising Effectiveness of Pre-Trained Vision Models for Control
Simone Parisi, Aravind Rajeswaran, Senthil Purushwalkam, and Abhinav Kumar Gupta. 2022 · 2022
Later among the works it cites.
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. In International Conference on Learning Representations (ICLR)
Ofir Press, Noah A. Smith, and Mike Lewis. 2022 · 2022
Later among the works it cites.
Real-World Robot Learning with Masked Visual Pre-training. In Conference on Robot Learning (CoRL)
Ilija Radosavovic, Tete Xiao, Stephen James, P. Abbeel, Jitendra Malik, and Trevor Darrell. 2022 · 2022
Later among the works it cites.
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley D. Edwards, Nicolas Manfred Otto Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, and Nando de Freitas. 2022 · 2022
Later among the works it cites.
Can Wikipedia Help Offline Reinforcement Learning?
Machel Reid, Yutaro Yamada, and Shixiang Shane Gu. 2022 · 2022
Later among the works it cites.
LAION-5B: An open large-scale dataset for training next generation image-text models. In Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS Datasets and Benchmarks)
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. 2022 · 2022
Later among the works it cites.
Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation. In Conference on Robot Learning (CoRL)
Mohit Shridhar, Lucas Manuelli, and Dieter Fox. 2022 · 2022
Later among the works it cites.
FLAVA: A Foundational Language And Vision Alignment Model. In Computer Vision and Pattern Recognition (CVPR) . 15617–15629
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. 2022 · 2022
Later among the works it cites.
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training. In Advances in Neural Information Processing Systems (NeurIPS)
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022 · 2022
Later among the works it cites.
Masked Visual Pre-training for Motor Control
Tete Xiao, Ilija Radosavovic, Trevor Darrell, and Jitendra Malik. 2022 · 2022
Later among the works it cites.
CoCa: Contrastive Captioners are Image-Text Foundation Models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. 2022 · 2022
Later among the works it cites.
Scaling Vision Transformers. In Computer Vision and Pattern Recognition (CVPR) . 1204–1213
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. 2022 · 2022
Later among the works it cites.
Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks. In International Conference on Learning Representations (ICLR)
Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi. 2023 · 2023
Closest in time.