Fetching the paper…
Reading the bibliography…
Chain-of-thought (CoT) reasoning has exhibited impressive performance in language models for solving complex tasks and answering questions.
UNIFIEDQA: Crossing Format Boundaries with a Single QA System
Khashabi, D.; Min, S.; Khot, T.; Sabharwal, A.; Tafjord, O.; Clark, P.; and Hajishirzi, H. 2020 · 1907
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Li, L. H.; Yatskar, M.; Yin, D.; Hsieh, C.-J.; and Chang, K.-W. 2019 · 1908
Earlier work this paper cites.
ROUGE: Recall-oriented understudy for gisting evaluation
Lin, C.-Y. 2003 · 2003
Earlier work this paper cites.
Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit
Bird, S.; Klein, E.; and Loper, E. 2009 · 2009
Earlier work this paper cites.
Generative Adversarial Nets
Goodfellow, I. J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A. C.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Kingma, D. P.; and Welling, M. 2014 · 2014
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Multi30K: Multilingual English-German Image Descriptions
Elliott, D.; Frank, S.; Sima’an, K.; and Specia, L. 2016 · 2016
Earlier work this paper cites.
LIUM-CVC Submissions for WMT17 Multimodal Translation Task
Caglayan, O.; Aransa, W.; Bardet, A.; García-Martínez, M.; Bougares, F.; Barrault, L.; Masana, M.; Herranz, L.; and van de Weijer, J. 2017 · 2017
Earlier work this paper cites.
Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Plummer, B. A.; Wang, L.; Cervantes, C. M.; Caicedo, J. C.; Hockenmaier, J.; and Lazebnik, S. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Earlier work this paper cites.
Bilinear attention networks
Kim, J.-H.; Jun, J.; and Zhang, B.-T. 2018 · 2018
Cited alongside, same era.
Dynamic fusion with intra-and inter-modality attention flow for visual question answering
Gao, P.; Jiang, Z.; You, H.; Lu, P.; Hoi, S. C.; Wang, X.; and Li, H. 2019 · 2019
Cited alongside, same era.
Deep modular co-attention networks for visual question answering
Yu, Z.; Yu, J.; Cui, Y.; Tao, D.; and Tian, Q. 2019 · 2019
Cited alongside, same era.
End-to-End Object Detection with Transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Cited alongside, same era.
Denoising Diffusion Probabilistic Models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Cited alongside, same era.
Dynamic context-guided capsule network for multimodal machine translation
Lin, H.; Meng, F.; Su, J.; Yin, Y.; Yang, Z.; Ge, Y.; Zhou, J.; and Luo, J. 2020 · 2020
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Later among the works it cites.
Denoising Diffusion Implicit Models
Song, J.; Meng, C.; and Ermon, S. 2021 · 2021
Later among the works it cites.
Wu, Z.; Kong, L.; Bi, W.; Li, X.; and Kao, B. 2021 · 2021
Later among the works it cites.
mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer
Xue, L.; Constant, N.; Roberts, A.; Kale, M.; Al-Rfou, R.; Siddhant, A.; Barua, A.; and Raffel, C. 2021 · 2021
Later among the works it cites.
Large Language Models are Zero-Shot Reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Cited alongside, same era.
A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine Translation
Yin, Y.; Meng, F.; Su, J.; Zhou, C.; Yang, Z.; Zhou, J.; and Luo, J. 2020 · 2020
Cited alongside, same era.
Neural Machine Translation with Universal Visual Representation
Zhang, Z.; Chen, K.; Wang, R.; Utiyama, M.; Sumita, E.; Li, Z.; and Zhao, H. 2020 · 2020
Cited alongside, same era.
Vilt: Vision-and-language transformer without convolution or region supervision
Kim, W.; Son, B.; and Kim, I. 2021 · 2021
Cited alongside, same era.
Generative Imagination Elevates Machine Translation
Long, Q.; Wang, M.; and Li, L. 2021 · 2021
Cited alongside, same era.
Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning
Lu, P.; Qiu, L.; Chen, J.; Xia, T.; Zhao, Y.; Zhang, W.; Yu, Z.; Liang, X.; and Zhu, S.-C. 2021 · 2021
Cited alongside, same era.
Diffusion Models already have a Semantic Latent Space
Kwon, M.; Jeong, J.; and Uh, Y. 2022 · 2022
Later among the works it cites.
On Vision Features in Multimodal Machine Translation
Li, B.; Lv, C.; Zhou, Z.; Zhou, T.; Xiao, T.; Ma, A.; and Zhu, J. 2022 · 2022
Later among the works it cites.
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.; Zhu, S.; Tafjord, O.; Clark, P.; and Kalyan, A. 2022 · 2022
Later among the works it cites.
Distill The Image to Nowhere: Inversion Knowledge Distillation for Multimodal Machine Translation
Peng, R.; Zeng, Y.; and Zhao, J. 2022 · 2022
Later among the works it cites.
High-Resolution Image Synthesis with Latent Diffusion Models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Later among the works it cites.
Chain of Thought Prompting Elicits Reasoning in Large Language Models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Chi, E. H.; Le, Q.; and Zhou, D. 2022 · 2022
Later among the works it cites.
Automatic Chain of Thought Prompting in Large Language Models
Zhang, Z.; Zhang, A.; Li, M.; and Smola, A. 2022 · 2022
Later among the works it cites.