Fetching the paper…
Reading the bibliography…
Image captioning models are typically trained by treating all samples equally, neglecting to account for mismatched or otherwise difficult data points.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019 · 1910
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
Self-paced learning for latent variable models
M Kumar, Benjamin Packer, and Daphne Koller. 2010 · 2010
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll’a r, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Automatic description generation from images: A survey of models, datasets, and evaluation measures
Raffaella Bernardi, Ruket Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, and Barbara Plank. 2016 · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
Room for improvement in automatic image description: an error analysis
Emiel van Miltenburg and Desmond Elliott. 2017 · 2017
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2018 · 2018
Earlier work this paper cites.
Inferring semantic layout for hierarchical text-to-image synthesis
Seunghoon Hong, Dingdong Yang, Jongwook Choi, and Honglak Lee. 2018 · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Text augmentation using bert for image captioning
Viktar Atliha and Dmitrij Šešok. 2020 · 2020
Cited alongside, same era.
A course-focused dual curriculum for image captioning
Mohammad Alsharid, Rasheed El-Bouri, Harshita Sharma, Lior Drukker, Aris T. Papageorghiou, and J. Alison Noble. 2021 · 2021
Cited alongside, same era.
Dual graph convolutional networks with transformer and curriculum learning for image captioning
Xinzhi Dong, Chengjiang Long, Wenju Xu, and Chunxia Xiao. 2021 · 2021
Cited alongside, same era.
CLIPScore: a reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Competence-based multimodal curriculum learning for medical report generation
Fenglin Liu, Shen Ge, and Xian Wu. 2021 · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
Cross-modal similarity-based curriculum learning for image captioning
Hongkuan Zhang, Saku Sugawara, Akiko Aizawa, Lei Zhou, Ryohei Sasano, and Koichi Takeda. 2022 · 2022
Later among the works it cites.
Learning from children: Improving image-caption pretraining via curriculum
Hammad Ayyubi, Rahul Lokesh, Alireza Zareian, Bo Wu, and Shih-Fu Chang. 2023 · 2023
Closest in time.
Synthetic data from diffusion models improves imagenet classification
Shekoofeh Azizi, Simon Kornblith, Chitwan Saharia, Mohammad Norouzi, and David J. Fleet. 2023 · 2023
Closest in time.
Synthcap: Augmenting transformers with synthetic data for image captioning
Davide Caffagni, Manuele Barraco, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2023 · 2023
Closest in time.
Extracting training data from diffusion models
Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramèr, Borja Balle, Daphne Ippolito, and Eric Wallace. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexander Quinn Nichol and Prafulla Dhariwal. 2021 · 2021
Cited alongside, same era.
Changing the world by changing the data
Anna Rogers. 2021 · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021 · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2021 · 2021
Cited alongside, same era.
Deduplicating training data mitigates privacy risks in language models
Nikhil Kandpal, Eric Wallace, and Colin Raffel. 2022 · 2022
Cited alongside, same era.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022 · 2022
Cited alongside, same era.
Few-shot learning with multilingual generative language models
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, et al. 2022 · 2022
Cited alongside, same era.
Closest in time.
Data curation alone can stabilize in-context learning
Ting-Yun Chang and Robin Jia. 2023 · 2023
Closest in time.
Lingua manga: A generic large language model centric system for data curation
Zui Chen, Lei Cao, and Sam Madden. 2023 · 2023
Closest in time.
Distilling model failures as directions in latent space
Saachi Jain, Hannah Lawrence, Ankur Moitra, and Aleksander Madry. 2023 · 2023
Closest in time.
Improving multimodal datasets with image captioning
Thao Nguyen, Samir Yitzhak Gadre, Gabriel Ilharco, Sewoong Oh, and Ludwig Schmidt. 2023 · 2023
Closest in time.
LMCap: Few-shot multilingual image captioning by retrieval augmented language model prompting
Rita Ramos, Bruno Martins, and Desmond Elliott. 2023 · 2023
Closest in time.
The curse of recursion: Training on generated data makes models forget
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson. 2023 · 2023
Closest in time.
Investigation finds ai image generation models trained on child abuse
David Thiel. 2023 · 2023
Closest in time.
Image as a foreign language: BEiT pretraining for vision and vision-language tasks
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, and Furu Wei. 2023 · 2023
Closest in time.
Multimodal data augmentation for image captioning using diffusion models
Changrong Xiao, Sean Xin Xu, and Kunpeng Zhang. 2023 · 2023
Closest in time.