Fetching the paper…
Reading the bibliography…
The current research direction in generative models, such as the recently developed GPT4, aims to find relevant knowledge information for multimodal and multilingual inputs to provide answers.
Rotate: Knowledge graph embedding by relational rotation in complex space
Sun, Z.; Deng, Z.-H.; Nie, J.-Y.; and Tang, J. 2019 · 1902
Earlier work this paper cites.
WordNet: a lexical database for English
Miller, G. A. 1995 · 1995
Earlier work this paper cites.
NLTK: the Natural Language Toolkit
Loper, E.; and Bird, S. 2002 · 2002
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
Auer, S.; Bizer, C.; Kobilarov, G.; Lehmann, J.; Cyganiak, R.; and Ives, Z. 2007 · 2007
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
A three-way model for collective learning on multi-relational data
Nickel, M.; Tresp, V.; Kriegel, H.-P.; et al. 2011 · 2011
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
Bordes, A.; Usunier, N.; Garcia-Duran, A.; Weston, J.; and Yakhnenko, O. 2013 · 2013
Earlier work this paper cites.
Knowledge graph embedding by translating on hyperplanes
Wang, Z.; Zhang, J.; Feng, J.; and Chen, Z. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia
Lehmann, J.; Isele, R.; Jakob, M.; Jentzsch, A.; Kontokostas, D.; Mendes, P. N.; Hellmann, S.; Morsey, M.; Van Kleef, P.; Auer, S.; et al. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Movieqa: Understanding stories in movies through question-answering
Tapaswi, M.; Zhu, Y.; Stiefelhagen, R.; Torralba, A.; Urtasun, R.; and Fidler, S. 2016 · 2016
Cited alongside, same era.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Cited alongside, same era.
Conceptnet 5.5: An open multilingual graph of general knowledge
Speer, R.; Chin, J.; and Havasi, C. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Kvqa: Knowledge-aware visual question answering
Shah, S.; Mishra, A.; Yadati, N.; and Talukdar, P. P. 2019 · 2019
Later among the works it cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Yu, Z.; Xu, D.; Yu, J.; Yu, T.; Zhao, Z.; Zhuang, Y.; and Tao, D. 2019 · 2019
Later among the works it cites.
Unsupervised Cross-lingual Representation Learning at Scale
Conneau, A.; Khandelwal, K.; Goyal, N.; Chaudhary, V.; Wenzek, G.; Guzmán, F.; Grave, É.; Ott, M.; Zettlemoyer, L.; and Stoyanov, V. 2020 · 2020
Later among the works it cites.
DramaQA: Character-centered video story understanding with hierarchical qa
Choi, S.; On, K.-W.; Heo, Y.-J.; Seo, A.; Jang, Y.; Lee, M.; and Zhang, B.-T. 2021 · 2021
Later among the works it cites.
Knowledge graph embedding for link prediction: A comparative analysis
Rossi, A.; Barbosa, D.; Firmani, D.; Matinata, A.; and Merialdo, P. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Convolutional 2d knowledge graph embeddings
Dettmers, T.; Minervini, P.; Stenetorp, P.; and Riedel, S. 2018 · 2018
Cited alongside, same era.
TVQA: Localized, Compositional Video Question Answering
Lei, J.; Yu, L.; Bansal, M.; and Berg, T. 2018 · 2018
Cited alongside, same era.
A Novel Embedding Model for Knowledge Base Completion Based on Convolutional Neural Network
Nguyen, T. D.; Nguyen, D. Q.; Phung, D.; et al. 2018 · 2018
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A.; and Manning, C. D. 2019 · 2019
Cited alongside, same era.
Explicit Knowledge-based Reasoning for Visual Question Answering
Wang, P.; Wu, Q.; Shen, C.; Dick, A.; and van den Hengel, A. 2017a
Cited in the paper.
Fvqa: Fact-based visual question answering
Wang, P.; Wu, Q.; Shen, C.; Dick, A.; and Van Den Hengel, A. 2017b
Cited in the paper.
Jimenez, C.; Russakovsky, O.; and Narasimhan, K. 2022 · 2022
Later among the works it cites.
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
Parcalabescu, L.; Cafagna, M.; Muradjan, L.; Frank, A.; Calixto, I.; and Gatt, A. 2022 · 2022
Later among the works it cites.
xGQA: Cross-Lingual Visual Question Answering
Pfeiffer, J.; Geigle, G.; Kamath, A.; Steitz, J.-M.; Roth, S.; Vulić, I.; and Gurevych, I. 2022 · 2022
Later among the works it cites.
A-okvqa: A benchmark for visual question answering using world knowledge
Schwenk, D.; Khandelwal, A.; Clark, C.; Marino, K.; and Mottaghi, R. 2022 · 2022
Later among the works it cites.
Complex embeddings for simple link prediction
Trouillon, T.; Welbl, J.; Riedel, S.; Gaussier, É.; and Bouchard, G. 2016 · 2080
Closest in time.