Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated impressive performance in natural language processing tasks by leveraging chain of thought (CoT) that enables step-by-step thinking.
Fast Graph Representation Learning with PyTorch Geometric
Fey, M.; and Lenssen, J. E. 2019 · 1903
Earlier work this paper cites.
Relational Graph Attention Networks
Busbridge, D.; Sherburn, D.; Cavallo, P.; and Hammerla, N. Y. 2019 · 1904
Earlier work this paper cites.
UNIFIEDQA: Crossing Format Boundaries with a Single QA System
Khashabi, D.; Min, S.; Khot, T.; Sabharwal, A.; Tafjord, O.; Clark, P.; and Hajishirzi, H. 2020 · 1907
Earlier work this paper cites.
VisualBERT: A Simple and Performant Baseline for Vision and Language
Li, L. H.; Yatskar, M.; Yin, D.; Hsieh, C.-J.; and Chang, K.-W. 2019 · 1908
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Semi-Supervised Classification with Graph Convolutional Networks
Kipf, T. N.; and Welling, M. 2017 · 2017
Earlier work this paper cites.
ConceptNet 5.5: An Open Multilingual Graph of General Knowledge
Speer, R.; Chin, J.; and Havasi, C. 2017 · 2017
Earlier work this paper cites.
Deeper Insights Into Graph Convolutional Networks for Semi-Supervised Learning
Li, Q.; Han, Z.; and Wu, X.-m. 2018 · 2018
Earlier work this paper cites.
KagNet: Knowledge-Aware Graph Networks for Commonsense Reasoning
Lin, B. Y.; Chen, X.; Chen, J.; and Ren, X. 2019 · 2019
Earlier work this paper cites.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Tan, H.; and Bansal, M. 2019 · 2019
Earlier work this paper cites.
End-to-End Object Detection with Transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Earlier work this paper cites.
Scalable Multi-Hop Relational Reasoning for Knowledge-Aware Question Answering
Feng, Y.; Chen, X.; Lin, B. Y.; Wang, P.; Yan, J.; and Ren, X. 2020 · 2020
Earlier work this paper cites.
What Does BERT with Vision Look At?
Li, L. H.; Yatskar, M.; Yin, D.; Hsieh, C.-J.; and Chang, K.-W. 2020 · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Cited alongside, same era.
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
Kim, W.; Son, B.; and Kim, I. 2021 · 2021
Cited alongside, same era.
IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
Lu, P.; Qiu, L.; Chen, J.; Xia, T.; Zhao, Y.; Zhang, W.; Yu, Z.; Liang, X.; and Zhu, S.-C. 2021 · 2021
Cited alongside, same era.
KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQA
Marino, K.; Chen, X.; Parikh, D.; Gupta, A.; and Rohrbach, M. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
Schwenk, D.; Khandelwal, A.; Clark, C.; Marino, K.; and Mottaghi, R. 2022 · 2022
Later among the works it cites.
GreaseLM: Graph REASoning Enhanced Language Models for Question Answering
Zhang, X.; Bosselut, A.; Yasunaga, M.; Ren, H.; Liang, P.; Manning, C. D.; and Leskovec, J. 2022 · 2022
Later among the works it cites.
Llama-adapter v2: Parameter-efficient visual instruction model
Gao, P.; Han, J.; Zhang, R.; Lin, Z.; Geng, S.; Zhou, A.; Zhang, W.; Lu, P.; He, C.; Yue, X.; et al. 2023 · 2023
Later among the works it cites.
SCITUNE: Aligning Large Language Models with Scientific Multimodal Instructions
Horawalavithana, S.; Munikoti, S.; Stewart, I.; and Kvinge, H. 2023 · 2023
Later among the works it cites.
Language Is Not All You Need: Aligning Perception with Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation
Wu, Z.; Kong, L.; Bi, W.; Li, X.; and Kao, B. 2021 · 2021
Cited alongside, same era.
QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering
Yasunaga, M.; Ren, H.; Bosselut, A.; Liang, P.; and Leskovec, J. 2021 · 2021
Cited alongside, same era.
Chen, W.; Ma, X.; Wang, X.; and Cohen, W. W. 2022 · 2022
Cited alongside, same era.
Scaling Instruction-Finetuned Language Models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; Webson, A.; Gu, S. S.; Dai, Z.; Suzgun, M.; Chen, X.; Chowdhery, A.; Castro-Ros, A.; Pellat, M.; Robinson, K.; Valter, D.; Narang, S.; Mishra, G.; Yu, A.; Zhao, V.; Huang, Y.; Dai, A.; Yu, H.; Petrov, S.; Chi, E. H.; Dean, J.; Devlin, J.; Roberts, A.; Zhou, D.; Le, Q. V.; and Wei, J. 2022 · 2022
Cited alongside, same era.
PromptCap: Prompt-Guided Task-Aware Image Captioning
Hu, Y.; Hua, H.; Yang, Z.; Shi, W.; Smith, N. A.; and Luo, J. 2022 · 2022
Cited alongside, same era.
Webly Supervised Concept Expansion for General Purpose Vision Models
Kamath, A.; Clark, C.; Gupta, T.; Kolve, E.; Hoiem, D.; and Kembhavi, A. 2022 · 2022
Cited alongside, same era.
On Vision Features in Multimodal Machine Translation
Li, B.; Lv, C.; Zhou, Z.; Zhou, T.; Xiao, T.; Ma, A.; and Zhu, J. 2022 · 2022
Cited alongside, same era.
Huang, S.; Dong, L.; Wang, W.; Hao, Y.; Singhal, S.; Ma, S.; Lv, T.; Cui, L.; Mohammed, O. K.; Liu, Q.; Aggarwal, K.; Chi, Z.; Bjorck, J.; Chaudhary, V.; Som, S.; Song, X.; and Wei, F. 2023 · 2023
Later among the works it cites.
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Later among the works it cites.
Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models
Luo, G.; Zhou, Y.; Ren, T.; Chen, S.; Sun, X.; and Ji, R. 2023 · 2023
Later among the works it cites.
ChatGPT, OpenAI
OpenAI. 2022 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Prompting Large Language Models with Answer Heuristics for Knowledge-based Visual Question Answering
Shao, Z.; Yu, Z.; Wang, M.; and Yu, J. 2023 · 2023
Later among the works it cites.
Wang, L.; Hu, Y.; He, J.; Xu, X.; Liu, N.; Liu, H.; and Shen, H. T. 2023 · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. 2023 · 2023
Later among the works it cites.
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Zhou, D.; Schärli, N.; Hou, L.; Wei, J.; Scales, N.; Wang, X.; Schuurmans, D.; Cui, C.; Bousquet, O.; Le, Q. V.; and Chi, E. H. 2023 · 2023
Later among the works it cites.
Neural machine translation: Challenges, progress and future
Zhang, J.; and Zong, C. 2020 · 2050
Closest in time.