Fetching the paper…
Reading the bibliography…
In this paper we explore the possibility of using OpenAI's CLIP to perform logically coherent grounded visual reasoning.
Macmillan
Daniel Kahneman (2011): Thinking, fast and slow · 2011
Earlier work this paper cites.
In: 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick & Devi Parikh (2015): VQA: Visual Question Answering · 2015
Earlier work this paper cites.
Available at https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser & Illia Polosukhin (2017): Attention is All you Need , pp. 5998–6008 · 2017
Earlier work this paper cites.
OpenAI blog
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever et al. (2019): Language models are unsupervised multitask learners · 2019
Earlier work this paper cites.
Available at https://vigilworkshop.github.io/static/papers/43.pdf
Sanjay Subramanian, Sameer Singh & Matt Gardner (2019): Analyzing Compositionality in Visual Question Answering · 2019
Cited alongside, same era.
In Andrea Vedaldi, Horst Bischof, Thomas Brox & Jan-Michael Frahm, editors: Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part V
Jae Sung Park, Chandra Bhagavatula, Roozbeh Mottaghi, Ali Farhadi & Yejin Choi (2020): VisualCOMET: Reasoning About the Dynamic Context of a Still Image · 2020
Cited alongside, same era.
In Danqi Chen, Jonathan Berant, Andrew McCallum & Sameer Singh, editors: 3rd Conference on Automated Knowledge Base Construction, AKBC 2021, Virtual, October 4-8, 2021
Chadi Helwe, Chloé Clavel & Fabian M. Suchanek (2021): Reasoning with Transformer-based Models: Deep Learning, but Shallow Reasoning · 2021
Cited alongside, same era.
In Marina Meila & Tong Zhang, editors: Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger & Ilya Sutskever (2021): Learning Transferable Visual Models From Natural Language Supervision · 2021
Cited alongside, same era.
Available at http://papers.nips.cc/paper_files/paper/2022/hash/960a172bc7fbf0177ccccbb411a7d800-Abstract-Conference.html
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman & Karén Simonyan (2022): Flamingo: a Visual Language Model for Few-Shot Learning · 2022
Later among the works it cites.
10.1109/CVPR52688.2022.01042
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser & Björn Ommer (2022): High-Resolution Image Synthesis with Latent Diffusion Models , pp. 10674–10685 · 2022
Later among the works it cites.
In Edward N. Zalta & Uri Nodelman, editors: The Stanford Encyclopedia of Philosophy
Zoltán Gendler Szabó (2022): Compositionality · 2022
Later among the works it cites.
MIT press, 10.7551/mitpress/2076.001.0001
Peter Gardenfors (2004): Conceptual spaces: The geometry of thought · 2076
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…