Fetching the paper…
Reading the bibliography…
Zero-shot referring expression comprehension aims at localizing bounding boxes in an image corresponding to provided textual prompts, which requires: (i) a fine-grained disentanglement of complex visual scene and textual context, and (ii) a capacity to understand relationships among disentangled entities.
Introduction to the conll-2005 shared task: Semantic role labeling
Xavier Carreras and Lluís Màrquez · 2005
Earlier work this paper cites.
Refining event extraction through cross-document inference
Heng Ji and Ralph Grishman · 2008
Earlier work this paper cites.
Nested named entity recognition
Jenny Rose Finkel and Christopher D Manning · 2009
Earlier work this paper cites.
Joint event extraction via structured prediction with global features
Qi Li, Heng Ji, and Liang Huang · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Hico: A benchmark for recognizing human-object interactions in images
Yu-Wei Chao, Zhan Wang, Yugeng He, Jiaxuan Wang, and Jia Deng · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Structured prediction energy networks
David Belanger and Andrew McCallum · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Deep semantic role labeling: What works and what’s next
Luheng He, Kenton Lee, Mike Lewis, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Scene graph generation from objects, phrases and region captions
Yikang Li, Wanli Ouyang, Bolei Zhou, Kun Wang, and Xiaogang Wang · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Scene graph generation by iterative message passing
Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei · 2017
Earlier work this paper cites.
Xu Han, Hao Zhu, Pengfei Yu, Ziyun Wang, Yuan Yao, Zhiyuan Liu, and Maosong Sun · 2018
Earlier work this paper cites.
Referring relationships
Ranjay Krishna, Ines Chami, Michael Bernstein, and Li Fei-Fei · 2018
Earlier work this paper cites.
Higher-order coreference resolution with coarse-to-fine inference
Kenton Lee, Luheng He, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
Earlier work this paper cites.
Neural motifs: Scene graph parsing with global context
Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi · 2018
Earlier work this paper cites.
A comprehensive survey of deep learning for image captioning
MD Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga · 2019
Earlier work this paper cites.
A unified mrc framework for named entity recognition
Xiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han, Fei Wu, and Jiwei Li · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi · 2019
Earlier work this paper cites.
Objects365: A large-scale, high-quality dataset for object detection
Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun · 2019
Cited alongside, same era.
Simple bert models for relation extraction and semantic role labeling
Peng Shi and Jimmy Lin · 2019
Cited alongside, same era.
Learning to compose dynamic tree structures for visual contexts
Kaihua Tang, Hanwang Zhang, Baoyuan Wu, Wenhan Luo, and Wei Liu · 2019
Cited alongside, same era.
Deep modular co-attention networks for visual question answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian · 2019
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Adapting clip for phrase localization without further training
Jiahao Li, Greg Shakhnarovich, and Raymond A Yeh · 2022
Later among the works it cites.
Gen-vlkt: Simplify association and enhance interaction understanding for hoi detection
Yue Liao, Aixi Zhang, Miao Lu, Yongliang Wang, Xiaobo Li, and Si Liu · 2022
Later among the works it cites.
Peft: State-of-the-art parameter-efficient fine-tuning methods
Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan · 2022
Later among the works it cites.
Flava: A foundational language and vision alignment model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela · 2022
Later among the works it cites.
From show to tell: A survey on deep learning-based image captioning
Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi, Silvia Cascianelli, Giuseppe Fiameni, and Rita Cucchiara · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen Gao, Jiarui Xu, Yuliang Zou, and Jia-Bin Huang · 2020
Cited alongside, same era.
Data annealing for informal language understanding tasks
Jing Gu and Zhou Yu · 2020
Cited alongside, same era.
Contrastive learning for weakly supervised phrase grounding
Tanmay Gupta, Arash Vahdat, Gal Chechik, Xiaodong Yang, Jan Kautz, and Derek Hoiem · 2020
Cited alongside, same era.
Consnet: Learning consistency graph for zero-shot human-object interaction detection
Ye Liu, Junsong Yuan, and Chang Wen Chen · 2020
Cited alongside, same era.
Grounded situation recognition
Sarah Pratt, Mark Yatskar, Luca Weihs, Ali Farhadi, and Aniruddha Kembhavi · 2020
Cited alongside, same era.
Nested named entity recognition via second-best sequence learning and decoding
Takashi Shibuya and Eduard Hovy · 2020
Cited alongside, same era.
Corefqa: Coreference resolution as query-based span prediction
Wei Wu, Fei Wang, Arianna Yuan, Fei Wu, and Jiwei Li · 2020
Cited alongside, same era.
Reclip: A strong zero-shot baseline for referring expression comprehension
Sanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner, Sameer Singh, and Anna Rohrbach · 2022
Later among the works it cites.
Panoptic scene graph generation
Jingkang Yang, Yi Zhe Ang, Zujin Guo, Kaiyang Zhou, Wayne Zhang, and Ziwei Liu · 2022
Later among the works it cites.
When and why vision-language models behave like bags-of-words, and what to do about it?
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou · 2022
Later among the works it cites.
Visual programming: Compositional visual reasoning without training
Tanmay Gupta and Aniruddha Kembhavi · 2023
Closest in time.
Incorporating structured representations into pretrained vision & language models using scene graphs
Roei Herzig, Alon Mendelson, Leonid Karlinsky, Assaf Arbelle, Rogerio Feris, Trevor Darrell, and Amir Globerson · 2023
Closest in time.
Structure-clip: Enhance multi-modal language representations with structure knowledge
Yufeng Huang, Jiji Tang, Zhuo Chen, Rongsheng Zhang, Xinfeng Zhang, Weijie Chen, Zeng Zhao, Tangjie Lv, Zhipeng Hu, and Wen Zhang · 2023
Closest in time.
Relational context learning for human-object interaction detection
Sanghyun Kim, Deunsol Jung, and Minsu Cho · 2023
Closest in time.
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Closest in time.
Uninext: Exploring a unified architecture for vision recognition
Fangjian Lin, Jianlong Yuan, Sitong Wu, Fan Wang, and Zhibin Wang · 2023
Closest in time.
Vgdiffzero: Text-to-image diffusion models can be zero-shot visual grounders
Xuyang Liu, Siteng Huang, Yachen Kang, Honggang Chen, and Donglin Wang · 2023
Closest in time.
Fgahoi: Fine-grained anchors for human-object interaction detection
Shuailei Ma, Yuefeng Wang, Shanze Wang, and Ying Wei · 2023
Closest in time.
Larger language models do in-context learning differently
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al · 2023
Closest in time.
Aligning bag of regions for open-vocabulary object detection
Size Wu, Wenwei Zhang, Sheng Jin, Wentao Liu, and Chen Change Loy · 2023
Closest in time.
Gentopia: A collaborative platform for tool-augmented llms
Binfeng Xu, Xukun Liu, Hua Shen, Zeyu Han, Yuhan Li, Murong Yue, Zhiyuan Peng, Yuchen Liu, Ziyu Yao, and Dongkuan Xu · 2023
Closest in time.
Zero-shot referring image segmentation with global-local context features
Seonghoon Yu, Paul Hongsuck Seo, and Jeany Son · 2023
Closest in time.
Diagnosing human-object interaction detectors
Fangrui Zhu, Yiming Xie, Weidi Xie, and Huaizu Jiang · 2023
Closest in time.
Muffin or chihuahua? challenging large vision-language models with multipanel vqa
Yue Fan, Jing Gu, Kaiwen Zhou, Qianqi Yan, Shan Jiang, Ching-Chen Kuo, Xinze Guan, and Xin Eric Wang · 2024
Closest in time.
Parameter-efficient fine-tuning for large models: A comprehensive survey
Zeyu Han, Chao Gao, Jinyang Liu, Sai Qian Zhang, et al · 2024
Closest in time.