Fetching the paper…
Reading the bibliography…
A handful of visual foundation models (VFMs) have recently emerged as the backbones for numerous downstream tasks.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2014
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
M. Everingham, S. M. A. Eslami, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2015
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, and Vincent Vanhoucke · 2016
Earlier work this paper cites.
Efficient knowledge distillation from an ensemble of teachers
Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Samuel Thomas, Jia Cui, and Bhuvana Ramabhadran · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Like what you like: Knowledge distill via neuron selectivity transfer
Zehao Huang and Naiyan Wang · 2017
Earlier work this paper cites.
Learning from multiple teacher networks
Shan You, Chang Xu, Chao Xu, and Dacheng Tao · 2017
Earlier work this paper cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2017
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
R. Cipolla, Y. Gal, and A. Kendall · 2018
Earlier work this paper cites.
Paraphrasing complex network: Network compression via factor transfer
Jangho Kim, SeongUk Park, and Nojun Kwak · 2018
Earlier work this paper cites.
Knowledge distillation by on-the-fly native ensemble, 2018
Xu Lan, Xiatian Zhu, and Shaogang Gong · 2018
Earlier work this paper cites.
Unsupervised feature learning via non-parametric instance discrimination
Zhirong Wu, Yuanjun Xiong, X Yu Stella, and Dahua Lin · 2018
Earlier work this paper cites.
Variational information distillation for knowledge transfer
S. Ahn, S. Hu, A. Damianou, N. D. Lawrence, and Z. Dai · 2019
Earlier work this paper cites.
Ensemble knowledge distillation for learning improved and efficient networks
Umar Asif, Jianbin Tang, and Stefan Harrer · 2019
Earlier work this paper cites.
A comprehensive overhaul of feature distillation
B. Heo, J. Kim, S. Yun, H. Park, N. Kwak, and J. Choi · 2019
Earlier work this paper cites.
Learning anytime predictions in neural networks via adaptive loss balancing
Hanzhang Hu, Debadeepta Dey, Martial Hebert, and J. Andrew Bagnell · 2019
Earlier work this paper cites.
GQA: a new dataset for compositional question answering over real-world images
Drew A. Hudson and Christopher D. Manning · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Improved knowledge distillation via teacher assistant
Seyed Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh · 2019
Earlier work this paper cites.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Cited alongside, same era.
Evalai: Towards better evaluation systems for ai agents, 2019
Deshraj Yadav, Rishabh Jain, Harsh Agrawal, Prithvijit Chattopadhyay, Taranjeet Singh, Akash Jain, Shiv Baran Singh, Stefan Lee, and Dhruv Batra · 2019
Cited alongside, same era.
MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark
MMSegmentation Contributors · 2020
Cited alongside, same era.
Adaptive multi-teacher multi-level knowledge distillation
Yuang Liu, Wei Zhang, and Jun Wang · 2020
Cited alongside, same era.
Feature-level ensemble knowledge distillation for aggregating knowledge from multiple networks
Seonguk Park and Nojun Kwak · 2020
Cited alongside, same era.
Designing network design spaces, 2020
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár · 2020
Metaformer is actually what you need for vision, 2022
Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan · 2022
Later among the works it cites.
Highlight every step: Knowledge distillation via collaborative teaching
Haoran Zhao, Xin Sun, Junyu Dong, Changrui Chen, and Zihe Dong · 2022
Later among the works it cites.
Foundational models defining a new era in vision: A survey and outlook, 2023
Muhammad Awais, Muzammal Naseer, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, and Fahad Shahbaz Khan · 2023
Closest in time.
Window attention is bugged: How not to interpolate position embeddings, 2023
Daniel Bolya, Chaitanya Ryali, Judy Hoffman, and Christoph Feichtenhofer · 2023
Closest in time.
Efficientvit: Multi-scale linear attention for high-resolution dense prediction, 2023
Han Cai, Junyan Li, Muyan Hu, Chuang Gan, and Song Han · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Self-training with noisy student improves imagenet classification
Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V. Le · 2020
Cited alongside, same era.
Model compression with two-stage multi-teacher knowledge distillation for web question answering system
Ze Yang, Linjun Shou, Ming Gong, Wutao Lin, and Daxin Jiang · 2020
Cited alongside, same era.
Reinforced multi-teacher selection for knowledge distillation, 2020
Fei Yuan, Linjun Shou, Jian Pei, Wutao Lin, Ming Gong, Yan Fu, and Daxin Jiang · 2020
Cited alongside, same era.
High-performance large-scale image recognition without normalization, 2021
Andrew Brock, Soham De, Samuel L. Smith, and Karen Simonyan · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers, 2021
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai · 2023
Closest in time.
Vision transformers need registers, 2023
Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski · 2023
Closest in time.
Data filtering networks, 2023
Alex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt, Alexander Toshev, and Vaishaal Shankar · 2023
Closest in time.
Datacomp: In search of the next generation of multimodal datasets, 2023
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Mussmann, Richard Vencu, Mehdi Cherti, Ranjay Krishna, Pang Wei Koh, Olga Saukh, Alexander Ratner, Shuran Song, Hannaneh Hajishirzi, Ali Farhadi, Romain Beaumont, Sewoong Oh, Alex Dimakis, Jenia Jitsev, Yair Carmon, Vaishaal Shankar, and Ludwig Schmidt · 2023
Closest in time.
Fastervit: Fast vision transformers with hierarchical attention, 2023
Ali Hatamizadeh, Greg Heinrich, Hongxu Yin, Andrew Tao, Jose M. Alvarez, Jan Kautz, and Pavlo Molchanov · 2023
Closest in time.
Navi: Category-agnostic image collections with high-quality 3d shape and pose annotations, 2023
Varun Jampani, Kevis-Kokitsi Maninis, Andreas Engelhardt, Arjun Karpur, Karen Truong, Kyle Sargent, Stefan Popov, André Araujo, Ricardo Martin-Brualla, Kaushal Patel, Daniel Vlasic, Vittorio Ferrari, Ameesh Makadia, Ce Liu, Yuanzhen Li, and Howard Zhou · 2023
Closest in time.
Ultralytics yolov8, 2023
Glenn Jocher, Ayush Chaurasia, and Jing Qiu · 2023
Closest in time.
Region-aware pretraining for open-vocabulary object detection with vision transformers, 2023
Dahun Kim, Anelia Angelova, and Weicheng Kuo · 2023
Closest in time.
Segment anything, 2023
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick · 2023
Closest in time.
Evaluating object hallucination in large vision-language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen · 2023
Closest in time.
Dinov2: Learning robust visual features without supervision, 2023
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski · 2023
Closest in time.
Sam-clip: Merging vision foundation models towards semantic and spatial understanding, 2023
Haoxiang Wang, Pavan Kumar Anasosalu Vasu, Fartash Faghri, Raviteja Vemulapalli, Mehrdad Farajtabar, Sachin Mehta, Mohammad Rastegari, Oncel Tuzel, and Hadi Pouransari · 2023
Closest in time.
Demystifying clip data
Hu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang, Russell Howes, Vasu Sharma, Shang-Wen Li, Gargi Ghosh, Luke Zettlemoyer, and Christoph Feichtenhofer · 2023
Closest in time.
Emernerf: Emergent spatial-temporal scene decomposition via self-supervision, 2023
Jiawei Yang, Boris Ivanovic, Or Litany, Xinshuo Weng, Seung Wook Kim, Boyi Li, Tong Che, Danfei Xu, Sanja Fidler, Marco Pavone, and Yue Wang · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Closest in time.
Multi-teacher knowledge distillation as an effective method for compressing ensembles of neural networks, 2023
Konrad Zuchniak · 2023
Closest in time.
Probing the 3D Awareness of Visual Foundation Models
Mohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis, Abhishek Kar, Yuanzhen Li, Michael Rubinstein, Deqing Sun, Leonidas Guibas, Justin Johnson, and Varun Jampani · 2024
Closest in time.
Denoising vision transformers, 2024
Jiawei Yang, Katie Z Luo, Jiefeng Li, Kilian Q Weinberger, Yonglong Tian, and Yue Wang · 2024
Closest in time.