Fetching the paper…
Reading the bibliography…
Evaluating the quality of automatically generated image descriptions is challenging, requiring metrics that capture various aspects such as grammaticality, coverage, correctness, and truthfulness.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
A technique for measuring attitude scale
RA Linkert. 1932 · 1932
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization . 65–72
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Re-evaluating the role of BLEU in machine translation research. In 11th conference of the european chapter of the association for computational linguistics . 249–256
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
Corpus-guided sentence generation of natural images. In Proceedings of the 2011 conference on empirical methods in natural language processing . 444–454
Yezhou Yang, Ching Teo, Hal Daumé III, and Yiannis Aloimonos. 2011 · 2011
Earlier work this paper cites.
Midge: Generating descriptions of images. In INLG 2012 Proceedings of the Seventh International Natural Language Generation Conference . 131–133
Margaret Mitchell, Xufeng Han, and Jeff Hayes. 2012 · 2012
Earlier work this paper cites.
Image description using visual dependency representations. In Proceedings of the 2013 conference on empirical methods in natural language processing . 1292–1302
Desmond Elliott and Frank Keller. 2013 · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013 · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions. In Proceedings of the IEEE international conference on computer vision . 433–440
Marcus Rohrbach, Wei Qiu, Ivan Titov, Stefan Thater, Manfred Pinkal, and Bernt Schiele. 2013 · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 . Springer, 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
See no evil, say no evil: Description generation from densely labeled images. In Proceedings of the Third Joint Conference on Lexical and Computational Semantics (* SEM 2014) . 110–120
Mark Yatskar, Michel Galley, Lucy Vanderwende, and Luke Zettlemoyer. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4566–4575
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3156–3164
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14 . Springer, 382–398
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition . 11–20
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. 2016 · 2016
Earlier work this paper cites.
Image captioning with deep bidirectional LSTMs. In Proceedings of the 24th ACM international conference on Multimedia . 988–997
Cheng Wang, Haojin Yang, Christian Bartz, and Christoph Meinel. 2016 · 2016
Earlier work this paper cites.
Robustness Analysis of Visual Question Answering Models by Basic Questions
Jia-Hong Huang. 2017 · 2017
Earlier work this paper cites.
VQABQ: Visual Question Answering by Basic Questions
Jia-Hong Huang, Modar Alfadly, and Bernard Ghanem. 2017 · 2017
Earlier work this paper cites.
On the automatic generation of medical imaging reports
Baoyu Jing, Pengtao Xie, and Eric Xing. 2017 · 2017
Earlier work this paper cites.
Areas of attention for image captioning. In Proceedings of the IEEE international conference on computer vision . 1242–1250
Marco Pedersoli, Thomas Lucas, Cordelia Schmid, and Jakob Verbeek. 2017 · 2017
Cited alongside, same era.
Shikhar Sharma, Layla El Asri, Hannes Schulz, and Jeremie Zumer. 2017 · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Robustness Analysis of Visual QA Models by Basic Questions
Longer Version for" Deep Context-Encoding Network for Retinal Image Captioning"
Jia-Hong Huang, Ting-Wei Wu, Chao-Han Huck Yang, and Marcel Worring. 2021d · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning . PMLR, 4904–4916
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision. In International Conference on Machine Learning
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jia-Hong Huang, Modar Alfadly, and Bernard Ghanem. 2018 · 2018
Cited alongside, same era.
Auto-classification of retinal diseases in the limit of sparse data using a two-streams machine learning model. In ACCV . Springer, 323–338
C-H Huck Yang, Fangyu Liu, Jia-Hong Huang, Meng Tian, I-Hung Lin, Yi Chieh Liu, Hiromasa Morikawa, Hao-Hsiang Yang, and Jesper Tegner. 2018 · 2018
Cited alongside, same era.
Textray: Mining clinical reports to gain a broad understanding of chest x-rays. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2018: 21st International Conference, Granada, Spain, September 16-20, 2018, Proceedings, Part II 11 . Springer, 553–561
Jonathan Laserson, Christine Dan Lantsman, Michal Cohen-Sfady, Itamar Tamir, Eli Goz, Chen Brestel, Shir Bar, Maya Atar, and Eldad Elnekave. 2018 · 2018
Cited alongside, same era.
Hybrid retrieval-generation reinforced agent for medical image report generation
Yuan Li, Xiaodan Liang, Zhiting Hu, and Eric P Xing. 2018 · 2018
Cited alongside, same era.
Synthesizing new retinal symptom images by multiple generative models. In ACCV . Springer, 235–250
Yi-Chieh Liu, Hao-Hsiang Yang, C-H Huck Yang, Jia-Hong Huang, Meng Tian, Hiromasa Morikawa, Yi-Chang James Tsai, and Jesper Tegner. 2018 · 2018
Cited alongside, same era.
A novel hybrid machine learning model for auto-classification of retinal diseases
C-H Huck Yang, Jia-Hong Huang, Fangyu Liu, Fang-Yi Chiu, Mengya Gao, Weifeng Lyu, Jesper Tegner, et al · 2018
Cited alongside, same era.
Silco: Show a few images, localize the common object. In ICCV . 5067–5076
Tao Hu, Pascal Mettes, Jia-Hong Huang, and Cees GM Snoek. 2019 · 2019
Cited alongside, same era.
Assessing the robustness of visual question answering
Jia-Hong Huang, Modar Alfadly, Bernard Ghanem, and Marcel Worring. 2019a · 2019
Cited alongside, same era.
The Dawn of Quantum Natural Language Processing
Riccardo Di Sipio, Jia-Hong Huang, Samuel Yen-Chi Chen, Stefano Mangini, and Marcel Worring. 2022 · 2022
Later among the works it cites.
Causal video summarizer for video exploration. In 2022 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 1–6
Jia-Hong Huang, Chao-Han Huck Yang, Pin-Yu Chen, Andrew Brown, and Marcel Worring. 2022b · 2022
Later among the works it cites.
High-Resolution Image Synthesis With Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Later among the works it cites.
Clair: Evaluating image captions with large language models
David Chan, Suzanne Petryk, Joseph E Gonzalez, Trevor Darrell, and John Canny. 2023 · 2023
Later among the works it cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. arXiv 2023
W Dai, J Li, D Li, AMH Tiong, J Zhao, W Wang, B Li, P Fung, and S Hoi. [n. d.] · 2023
Later among the works it cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Albert Li, Pascale Fung, and Steven C. H. Hoi. 2023 · 2023
Later among the works it cites.
Eva: Exploring the limits of masked visual representation learning at scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 19358–19369
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. 2023 · 2023
Later among the works it cites.
Jia-Hong Huang, Modar Alfadly, Bernard Ghanem, and Marcel Worring. 2023a · 2023
Later among the works it cites.
Query-Based Video Summarization with Pseudo Label Supervision. In 2023 IEEE International Conference on Image Processing (ICIP) . IEEE, 1430–1434
Jia-Hong Huang, Luka Murn, Marta Mrak, and Marcel Worring. 2023b · 2023
Later among the works it cites.
Conditional Modeling Based Automatic Video Summarization
Jia-Hong Huang, Chao-Han Huck Yang, Pin-Yu Chen, Min-Hung Chen, and Marcel Worring. 2023d · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
A brief overview of ChatGPT: The history, status quo and potential future development
Tianyu Wu, Shizhu He, Jingping Liu, Siqi Sun, Kang Liu, Qing-Long Han, and Yang Tang. 2023a · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Haotong Zhang, Joseph Gonzalez, and Ion Stoica. 2023 · 2023
Later among the works it cites.
Multi-modal Video Summarization. In Proceedings of the 2024 International Conference on Multimedia Retrieval . 1214–1218
Jia-Hong Huang. 2024 · 2024
Closest in time.
Optimizing Numerical Estimation and Operational Efficiency in the Legal Domain through Large Language Models. In ACM International Conference on Information and Knowledge Management (CIKM)
Jia-Hong Huang, Chao-Chun Yang, Yixian Shen, Alessio M Pacces, and Evangelos Kanoulas. 2024 · 2024
Closest in time.
Ada-HGNN: Adaptive Sampling for Scalable Hypergraph Neural Networks
Shuai Wang, David W Zhang, Jia-Hong Huang, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, and Marcel Worring. 2024 · 2024
Closest in time.
Towards Fine-Grained Citation Evaluation in Generated Text: A Comparative Analysis of Faithfulness Metrics. In Proceedings of the 2024 International Natural Language Generation Conference (INLG)
Weijia Zhang, Mohammad Aliannejadi, Yifei Yuan, Jiahuan Pei, Jia-Hong Huang, and Evangelos Kanoulas. 2024b · 2024
Closest in time.
QFMTS: Generating Query-Focused Summaries over Multi-Table Inputs. In Proceedings of the 2024 European Conference on Artificial Intelligence (ECAI)
Weijia Zhang, Vaishali Pal, Jia-Hong Huang, Evangelos Kanoulas, and Maarten de Rijke. 2024d · 2024
Closest in time.
Enhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models. In Proceedings of the 2024 International Conference on Multimedia Retrieval . 978–987
Hongyi Zhu, Jia-Hong Huang, Stevan Rudinac, and Evangelos Kanoulas. 2024 · 2024
Closest in time.