Fetching the paper…
Reading the bibliography…
Large Vision Language Models (VLMs) are now the de facto state-of-the-art for a number of tasks including visual question answering, recognising objects, and spatial referral.
Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., Serre, T.: HMDB: a large video database for human motion recognition. In: International Conference on Computer Vision (ICCV) (2011)
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
Plummer, B.A., Wang, L., Cervantes, C.M., Caicedo, J.C., Hockenmaier, J., Lazebnik, S.: Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models. In: International Conference on Computer Vision (ICCV) (2015)
2015
Earlier work this paper cites.
Damen, D., Doughty, H., Farinella, G.M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., Wray, M.: Scaling Egocentric Vision: The EPIC-KITCHENS Dataset. In: European Conference on Computer Vision (ECCV) (2018)
2018
Earlier work this paper cites.
Li, Y., Liu, M., Rehg, J.M.: In the Eye of Beholder: Joint Learning of Gaze and Actions in First Person Video. In: European Conference on Computer Vision (ECCV) (2018)
2018
Earlier work this paper cites.
Sigurdsson, G.A., Gupta, A., Schmid, C., Farhadi, A., Alahari, K.: Actor and Observer: Joint Modeling of First and Third-Person Videos. In: Computer Vision and Pattern Recognition (CVPR) (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Narasimhaswamy, S., Wei, Z., Wang, Y., Zhang, J., Hoai, M.: Contextual Attention for Hand Detection in the Wild. In: International Conference on Computer Vision (ICCV) (2019)
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language models are few-shot learners. Advances in Neural Information Processing Systems (NeurIPS) (2020)
2020
Earlier work this paper cites.
Narasimhaswamy, S., Nguyen, T., Nguyen, M.H.: Detecting Hands and Recognizing Physical Contact in the Wild. Advances in Neural Information Processing Systems (NeurIPS) (2020)
2020
Earlier work this paper cites.
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research (JMLR) (2020)
2020
Earlier work this paper cites.
Shan, D., Geng, J., Shu, M., Fouhey, D.F.: Understanding Human Hands in Contact at Internet Scale. In: Computer Vision and Pattern Recognition (CVPR) (2020)
2020
Earlier work this paper cites.
Kamath, A., Singh, M., LeCun, Y., Synnaeve, G., Misra, I., Carion, N.: MDETR – Modulated Detection for End-to-End Multi-Modal Understanding. In: International Conference on Computer Vision (ICCV) (2021)
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning Transferable Visual Models From Natural Language Supervision. In: International Conference on Machine Learning (ICML) (2021)
2021
Earlier work this paper cites.
Alayrac, J.B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al.: Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems (NeurIPS) (2022)
2022
Earlier work this paper cites.
Bansal, S., Arora, C., Jawahar, C.: My View is the Best View: Procedure Learning from Egocentric Videos. In: European Conference on Computer Vision (ECCV) (2022)
2022
Earlier work this paper cites.
Damen, D., Doughty, H., Farinella, G.M., Furnari, A., Ma, J., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., Wray, M.: Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100. International Journal of Computer Vision (IJCV) (2022)
2022
Earlier work this paper cites.
Darkhalil, A., Shan, D., Zhu, B., Ma, J., Kar, A., Higgins, R., Fidler, S., Fouhey, D., Damen, D.: EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations. In: Proceedings of the Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks (2022)
2022
Earlier work this paper cites.
Dou, Z.Y., Kamath, A., Gan, Z., Zhang, P., Wang, J., Li, L., Liu, Z., Liu, C., LeCun, Y., Peng, N., et al.: Coarse-to-fine vision-language pre-training with fusion in the backbone. Advances in Neural Information Processing Systems (NeurIPS) (2022)
2022
Earlier work this paper cites.
Grauman, K., et al.: Ego4D: Around the World in 3,000 Hours of Egocentric Video. In: Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Earlier work this paper cites.
Hu, E.J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: LoRA: Low-Rank Adaptation of Large Language Models. In: International Conference on Learning Representations (ICLR) (2022)
2022
Earlier work this paper cites.
Li, J., Li, D., Xiong, C., Hoi, S.: BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation. In: International Conference on Machine Learning (ICML) (2022)
2022
Cited alongside, same era.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al.: Training language models to follow instructions with human feedback. Neural Information Processing Systems (NeurIPS) (2022)
2022
Cited alongside, same era.
Qinghong Lin, K., Jinpeng Wang, A., Soldan, M., Wray, M., Yan, R., Zhongcong Xu, E., Gao, D., Tu, R., Zhao, W., Kong, W., et al.: Egocentric Video-Language Pretraining. In: Advances in Neural Information Processing Systems (NeurIPS) (2022)
2022
Cited alongside, same era.
Sener, F., Chatterjee, D., Shelepov, D., He, K., Singhania, D., Wang, R., Yao, A.: Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities. Computer Vision and Patter Recognition (CVPR) (2022)
Hazra, R., Chen, B., Rai, A., Kamra, N., Desai, R.: EgoTV: Egocentric Task Verification from Natural Language Task Descriptions. In: International Conference on Computer Vision (ICCV) (2023)
2023
Later among the works it cites.
Kurita, S., Katsura, N., Onami, E.: RefEgo: Referring Expression Comprehension Dataset from First-Person Perception of Ego4D. In: International Conference on Computer Vision (ICCV) (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual Instruction Tuning. Advances in Neural Information Processing Systems (NeurIPS) (2023)
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Pramanick, S., Song, Y., Nag, S., Lin, K.Q., Shah, H., Shou, M.Z., Chellappa, R., Zhang, P.: EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone. In: International Conference on Computer Vision (ICCV) (2023)
2023
Later among the works it cites.
Ragusa, F., Furnari, A., Farinella, G.M.: MECCANO: A Multimodal Egocentric Dataset for Humans Behavior Understanding in the Industrial-like Domain. Computer Vision and Image Understanding (CVIU) (2023)
2023
Later among the works it cites.
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., Zhuang, Y.: HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. Advances in Neural Information Processing Systems (NeurIPS) (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Zhao, Y., Misra, I., Krähenbühl, P., Girdhar, R.: Learning Video Representations from Large Language Models. In: Computer Vision and Pattern Recognition (CVPR) (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
Guo, Q., De Mello, S., Yin, H., Byeon, W., Chun Cheung, K., Yu, Y., Luo, P., Liu, S.: Regiongpt. Computer Vision and Pattern Recognition (CVPR) (2024)
2024
Closest in time.
Journal, W.S.: I Spent 24 Hours Wearing Apple’s Vision Pro Headset | WSJ (2024), https://youtu.be/8xI10SFgzQ8?t=283 , Last accessed: 2024-03-04
2024
Closest in time.