Fetching the paper…
Reading the bibliography…
Vision and Language (VL) models offer an effective method for aligning representation spaces of images and text, leading to numerous applications such as cross-modal retrieval, visual question answering, captioning, and more.
“Learning multiple layers of features from tiny images”
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
“Similarity Constrained Latent Support Vector Machine: An Application to Weakly Supervised Action Classification”
Nataliya Shapovalova, Arash Vahdat, Kevin. Cannons, Tian Lan and Greg Mori · 2012
Earlier work this paper cites.
“Finding Actors and Actions in Movies”
Piotr Bojanowski, Francis. Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid and Josef Sivic · 2013
Earlier work this paper cites.
“Microsoft coco: Common objects in context”
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár and C Zitnick · 2014
Earlier work this paper cites.
“Is object localization for free? - Weakly-supervised learning with convolutional neural networks”
Maxime Oquab, on Bottou, Ivan Laptev and Josef Sivic · 2015
Earlier work this paper cites.
“Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models”
Bryan Plummer, Liwei Wang, Chris Cervantes, Juan Caicedo, Julia Hockenmaier and Svetlana Lazebnik · 2015
Earlier work this paper cites.
“Locality Sensitive Deep Learning for Detection and Classification of Nuclei in Routine Colon Cancer Histology Images”
Korsuk Sirinukunwattana, Shan-e-Ahmed Raza, Yee-Wah Tsang, David.. Snead, Ian. Cree and Nasir. Rajpoot · 2016
Earlier work this paper cites.
“Learning Deep Features for Discriminative Localization”
Bolei Zhou, Aditya Khosla, gata Lapedriza, Aude Oliva and Antonio Torralba · 2016
Earlier work this paper cites.
“Scene Graph Generation by Iterative Message Passing”
Danfei Xu, Yuke Zhu, Christopher. Choy and Li Fei-Fei · 2017
Earlier work this paper cites.
“Learning from Video and Text via Large-Scale Discriminative Clustering”
Antoine Miech, Jean-Baptiste Alayrac, Piotr Bojanowski, Ivan Laptev and Josef Sivic · 2017
Earlier work this paper cites.
“Multiple-Instance Learning for Medical Image and Video Analysis”
Gwénolé Quellec, Guy Cazuguel, Béatrice Cochener and Mathieu Lamard · 2017
Earlier work this paper cites.
“Visual genome: Connecting language and vision using crowdsourced dense image annotations”
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li and David Shamma · 2017
Earlier work this paper cites.
“Mapping Images to Scene Graphs with Permutation-Invariant Structured Prediction”
Roei Herzig, Moshiko Raboh, Gal Chechik, Jonathan Berant and Amir Globerson · 2018
Earlier work this paper cites.
“Referring Relationships”
Ranjay Krishna, Ines Chami, Michael. Bernstein and Li Fei-Fei · 2018
Earlier work this paper cites.
“Object Level Visual Reasoning in Videos”
Fabien Baradel, Natalia Neverova, Christian Wolf, Julien Mille and Greg Mori · 2018
Earlier work this paper cites.
“Relational inductive biases, deep learning, and graph networks”
Peter Battaglia, Jessica Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro and Ryan Faulkner · 2018
Earlier work this paper cites.
“Compositional Learning for Human Object Interaction”
Keizo Kato, Yin Li and Abhinav Gupta · 2018
Earlier work this paper cites.
“Videos as Space-Time Region Graphs”
Xiaolong Wang and Abhinav Gupta · 2018
Earlier work this paper cites.
“Image generation from scene graphs”
Justin Johnson, Agrim Gupta and Li Fei-Fei · 2018
Earlier work this paper cites.
“A flexible model for training action localization with varying levels of supervision”
Guilhem Chéron, Jean-Baptiste Alayrac, Ivan Laptev and Cordelia Schmid · 2018
Earlier work this paper cites.
“Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning”
Piyush Sharma, Nan Ding, Sebastian Goodman and Radu Soricut · 2018
Earlier work this paper cites.
Patrick Helber, Benjamin Bischke, Andreas Dengel and Damian Borth · 2018
Earlier work this paper cites.
“VisualBERT: A Simple and Performant Baseline for Vision and Language”
Liunian Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh and Kai-Wei Chang · 2019
Earlier work this paper cites.
“LXMERT: Learning Cross-Modality Encoder Representations from Transformers”
Hao Tan and Mohit Bansal · 2019
Earlier work this paper cites.
“Spatio-temporal action graph networks”
Roei Herzig, Elad Levi, Huijuan Xu, Hang Gao, Eli Brosh, Xiaolong Wang, Amir Globerson and Trevor Darrell · 2019
Earlier work this paper cites.
“Action Genome: Actions as Composition of Spatio-temporal Scene Graphs”
Jingwei Ji, Ranjay Krishna, Li Fei-Fei and Juan Niebles · 2019
Earlier work this paper cites.
“Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection”
Keren Ye, Mingda Zhang, Adriana Kovashka, Wei Li, Danfeng Qin and Jesse Berent · 2019
Earlier work this paper cites.
“Hake: Human activity knowledge engine”
Yong-Lu Li, Liang Xu, Xinpeng Liu, Xijie Huang, Yue Xu, Mingyang Chen, Ze Ma, Shiyi Wang, Hao-Shu Fang and Cewu Lu · 2019
Earlier work this paper cites.
“UNITER: UNiversal Image-TExt Representation Learning”
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng and Jingjing Liu · 2020
Earlier work this paper cites.
“Oscar: Object-semantics aligned pre-training for vision-language tasks”
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong and Furu Wei · 2020
Earlier work this paper cites.
“Learning Object Detection from Captions via Textual Scene Attributes”
Achiya Jerbi, Roei Herzig, Jonathan Berant, Gal Chechik and Amir Globerson · 2020
Cited alongside, same era.
“Differentiable Scene Graphs”
Moshiko Raboh, Roei Herzig, Gal Chechik, Jonathan Berant and Amir Globerson · 2020
Cited alongside, same era.
“DRG: Dual Relation Graph for Human-Object Interaction Detection”
Chen Gao, Jiarui Xu, Yuliang Zou and Jia-Bin Huang · 2020
Cited alongside, same era.
“Something-Else: Compositional Action Recognition with Spatial-Temporal Interaction Networks”
Joanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu, Xiaolong Wang and Trevor Darrell · 2020
Cited alongside, same era.
“Learning Canonical Representations for Scene Graph to Image Generation”
Roei Herzig, Amir Bar, Huijuan Xu, Gal Chechik, Trevor Darrell and Amir Globerson · 2020
Cited alongside, same era.
“Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality”
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela and Candace Ross · 2022
Later among the works it cites.
“Teaching Structured Vision&Language Concepts to Vision&Language Models”
Sivan Doveh, Assaf Arbelle, Sivan Harary, Rameswar Panda, Roei Herzig, Eli Schwartz, Donghyun Kim, Raja Giryes, Rogerio Feris and Shimon Ullman · 2022
Later among the works it cites.
“Opt: Open pre-trained transformer language models”
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li and Xi Lin · 2022
Later among the works it cites.
“Vision-Language Pre-Training with Triple Contrastive Learning”
Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi and Junzhou Huang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“End-to-end learning of visual representations from uncurated instructional videos”
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic and Andrew Zisserman · 2020
Cited alongside, same era.
“Grounded situation recognition”
Sarah Pratt, Mark Yatskar, Luca Weihs, Ali Farhadi and Aniruddha Kembhavi · 2020
Cited alongside, same era.
“Learning transferable visual models from natural language supervision”
Alec Radford, Jong Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin and Jack Clark · 2021
Cited alongside, same era.
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc. Le, Yunhsuan Sung, Zhen Li and Tom Duerig · 2021
Cited alongside, same era.
“Open-vocabulary object detection using captions”
Alireza Zareian, Kevin Rosa, Derek Hu and Shih-Fu Chang · 2021
Cited alongside, same era.
“Open-vocabulary object detection via vision and language knowledge distillation”
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo and Yin Cui · 2021
Cited alongside, same era.
“Laion-400m: Open dataset of clip-filtered 400 million image-text pairs”
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev and Aran Komatsuzaki · 2021
Cited alongside, same era.
Shashank Goel, Hritik Bansal, Sumit Bhatia, Ryan Rossi, Vishwa Vinay and Aditya Grover · 2022
Later among the works it cites.
“PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining”
Yuting Gao, Jinfeng Liu, Zihan Xu, Jun Zhang, Ke Li and Chunhua Shen · 2022
Later among the works it cites.
“Bringing Image Scene Structure to Video via Frame-Clip Consistency of Object Tokens”
Elad Avraham, Roei Herzig, Karttikeya Mangalam, Amir Bar, Anna Rohrbach, Leonid Karlinsky, Trevor Darrell and Amir Globerson · 2022
Later among the works it cites.
“Object-Region Video Transformers”
Roei Herzig, Elad Ben-Avraham, Karttikeya Mangalam, Amir Bar, Gal Chechik, Anna Rohrbach, Trevor Darrell and Amir Globerson · 2022
Later among the works it cites.
“FETA: Towards Specializing Foundation Models for Expert Task Applications”
Amit Alfassy, Assaf Arbelle, Oshri Halimi, Sivan Harary, Roei Herzig, Eli Schwartz, Rameswar Panda, Michele Dolfi, Christoph Auer and Kate Saenko · 2022
Later among the works it cites.
“Is a caption worth a thousand images? a controlled study for representation learning”
Shibani Santurkar, Yann Dubois, Rohan Taori, Percy Liang and Tatsunori Hashimoto · 2022
Later among the works it cites.
“ConStruct-VL: Data-Free Continual Structured VL Concepts Learning”
James Smith, Paola Cascante-Bonilla, Assaf Arbelle, Donghyun Kim, Rameswar Panda, David Cox, Diyi Yang, Zsolt Kira, Rogerio Feris and Leonid Karlinsky · 2022
Later among the works it cites.
“Lit: Zero-shot transfer with locked-image text tuning”
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov and Lucas Beyer · 2022
Later among the works it cites.
“ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models”
Chunyuan Li, Haotian Liu, Liunian Li, Pengchuan Zhang, Jyoti Aneja, Jianwei Yang, Ping Jin, Yong Lee, Houdong Hu, Zicheng Liu and Jianfeng Gao · 2022
Later among the works it cites.
“GPT-3.5”, GitHub repository, 2023
OpenAI · 2023
Closest in time.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Closest in time.
“Llama: Open and efficient foundation language models”
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro and Faisal Azhar · 2023
Closest in time.
“Stanford Alpaca: An Instruction-following LLaMA model”
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang and Tatsunori. Hashimoto · 2023
Closest in time.
“Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality”, 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph. Gonzalez, Ion Stoica and Eric. Xing · 2023
Closest in time.
“BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models”
Junnan Li, Dongxu Li, Silvio Savarese and Steven Hoi · 2023
Closest in time.
“Minigpt-4: Enhancing vision-language understanding with advanced large language models”
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li and Mohamed Elhoseiny · 2023
Closest in time.
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Closest in time.
“When and why Vision-Language Models behave like Bags-of-Words, and what to do about it?”
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky and James Zou · 2023
Closest in time.
“Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond”, 2023
Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Bing Yin and Xia Hu · 2023
Closest in time.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander Berg and Wan-Yen Lo · 2023
Closest in time.
Roei Herzig, Alon Mendelson, Leonid Karlinsky, Assaf Arbelle, Rogerio Feris, Trevor Darrell and Amir Globerson · 2023
Closest in time.
“Visual Spatial Reasoning”
Fangyu Liu, Guy Emerson and Nigel Collier · 2023
Closest in time.
Roei Herzig, Ofir Abramovich, Elad Ben-Avraham, Assaf Arbelle, Leonid Karlinsky, Ariel Shamir, Trevor Darrell and Amir Globerson · 2023
Closest in time.
“Learning to Detect Human-Object Interactions With Knowledge”
Bingjie Xu, Yongkang Wong, Junnan Li, Qi Zhao and M. Kankanhalli · 2028
Closest in time.
“Handling label noise in video classification via multiple instance learning”
Thomas Leung, Yang Song and John. Zhang · 2063
Closest in time.