Fetching the paper…
Reading the bibliography…
Large-scale, pre-trained neural networks have demonstrated strong capabilities in various tasks, including zero-shot image segmentation.
Pengi: An implementation of a theory of activity
Agre, P. E. and Chapman, D · 1987
Earlier work this paper cites.
The thing that we tried didn’t work very well: Deictic representation in reinforcement learning
Finney, S., Gardiol, N., Kaelbling, L. P., and Oates, T · 2002
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Russell, S. and Norvig, P · 2010
Earlier work this paper cites.
ReferItGame: Referring to objects in photographs of natural scenes
Kazemzadeh, S., Ordonez, V., Matten, M., and Berg, T · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Visual relationship detection with language priors
Lu, C., Krishna, R., Bernstein, M., and Fei-Fei, L · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C. L., and Girshick, R · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al · 2017
Earlier work this paper cites.
End-to-end differentiable proving
Rocktäschel, T. and Riedel, S · 2017
Earlier work this paper cites.
Scene graph generation by iterative message passing
Xu, D., Zhu, Y., Choy, C. B., and Fei-Fei, L · 2017
Earlier work this paper cites.
Learning explanatory rules from noisy data
Evans, R. and Grefenstette, E · 2018
Earlier work this paper cites.
A review of semantic segmentation using deep neural networks
Guo, Y., Liu, Y., Georgiou, T., and Lew, M. S · 2018
Earlier work this paper cites.
Using binary decision diagrams to enumerate inductive logic programming solutions
Shindo, H., Nishino, M., and Yamamoto, A · 2018
Earlier work this paper cites.
Understanding convolution for semantic segmentation
Wang, P., Chen, P., Yuan, Y., Liu, D., Huang, Z., Hou, X., and Cottrell, G. W · 2018
Earlier work this paper cites.
Neural-symbolic VQA: disentangling reasoning from vision and language understanding
Yi, K., Wu, J., Gan, C., Torralba, A., Kohli, P., and Tenenbaum, J · 2018
Earlier work this paper cites.
Neural motifs: Scene graph parsing with global context
Zellers, R., Yatskar, M., Thomson, S., and Choi, Y · 2018
Earlier work this paper cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
Mao, J., Gan, C., Kohli, P., Tenenbaum, J. B., and Wu, J · 2019
Earlier work this paper cites.
LXMERT: learning cross-modality encoder representations from transformers
Tan, H. and Bansal, M · 2019
Earlier work this paper cites.
Learning to compose dynamic tree structures for visual contexts
Tang, K., Zhang, H., Wu, B., Luo, W., and Liu, W · 2019
Cited alongside, same era.
Neuro-symbolic visual reasoning: Disentangling "Visual" from "Reasoning"
Amizadeh, S., Palangi, H., Polozov, A., Huang, Y., and Koishida, K · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
GPS-Net: Graph property sensing network for scene graph generation
Lin, X., Ding, C., Zeng, J., and Tao, D · 2020
Cited alongside, same era.
Object-centric learning with slot attention
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2020
Cited alongside, same era.
Leveraging large language models to generate answer set programs
Ishay, A., Yang, Z., and Lee, J · 2023
Later among the works it cites.
Segment anything
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollar, P., and Girshick, R · 2023
Later among the works it cites.
Lisa: Reasoning segmentation via large language model
Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., and Jia, J · 2023
Later among the works it cites.
Not all neuro-symbolic concepts are created equal: Analysis and mitigation of reasoning shortcuts
Marconato, E., Teso, S., Vergari, A., and Passerini, A · 2023
Later among the works it cites.
Neural-symbolic predicate invention: Learning relational concepts from visual scenes
Sha, J., Shindo, H., Kersting, K., and Dhami, D. S · 2023
Later among the works it cites.
Large language models can be easily distracted by irrelevant context
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning reasoning strategies in end-to-end differentiable proving
Minervini, P., Riedel, S., Stenetorp, P., Grefenstette, E., and Rocktäschel, T · 2020
Cited alongside, same era.
Clevrer: Collision events for video representation and reasoning
Yi, K., Gan, C., Li, Y., Kohli, P., Wu, J., Torralba, A., and Tenenbaum, J. B · 2020
Cited alongside, same era.
Context-aware scene graph generation with Seq2Seq transformers
Lu, Y., Rai, H., Chang, J., Knyazev, B., Yu, G., Shekhar, S., Taylor, G. W., and Volkovs, M · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Neuro-symbolic forward reasoning
Shindo, H., Dhami, D. S., and Kersting, K · 2021
Cited alongside, same era.
Right for the right concept: Revising neuro-symbolic concepts by interacting with their explanations
Stammer, W., Schramowski, P., and Kersting, K · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J. L., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Bińkowski, M. a., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Cited alongside, same era.
Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., Schärli, N., and Zhou, D · 2023
Later among the works it cites.
α \alpha ilp: thinking visual scenes as differentiable logic programs
Shindo, H., Pfanschilling, V., Dhami, D. S., and Kersting, K · 2023
Later among the works it cites.
Vision relation transformer for unbiased scene graph generation
Sudhakaran, G., Dhami, D. S., Kersting, K., and Roth, S · 2023
Later among the works it cites.
Vipergpt: Visual inference via python execution for reasoning
Surís, D., Menon, S., and Vondrick, C · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
From word models to world models: Translating from natural language to the probabilistic language of thought
Wong, L., Grand, G., Lew, A. K., Goodman, N. D., Mansinghka, V. K., Andreas, J., and Tenenbaum, J. B · 2023
Later among the works it cites.
Coupling large language models with logic programming for robust and general reasoning from text
Yang, Z., Ishay, A., and Lee, J · 2023
Later among the works it cites.
Differentiable logic machines
Zimmer, M., Feng, X., Glanois, C., JIANG, Z., Zhang, J., Weng, P., Li, D., HAO, J., and Liu, W · 2023
Later among the works it cites.
Segment everything everywhere all at once
Zou, X., Yang, J., Zhang, H., Li, F., Li, L., Wang, J., Wang, L., Gao, J., and Lee, Y. J · 2023
Later among the works it cites.
Large language models cannot self-correct reasoning yet
Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., and Zhou, D · 2024
Closest in time.
Grounded SAM: assembling open-world models for diverse visual tasks
Ren, T., Liu, S., Zeng, A., Lin, J., Li, K., Cao, H., Chen, J., Huang, X., Chen, Y., Yan, F., Zeng, Z., Zhang, H., Li, F., Yang, J., Li, H., Jiang, Q., and Zhang, L · 2024
Closest in time.
EXPIL: explanatory predicate invention for learning in games
Sha, J., Shindo, H., Delfosse, Q., Kersting, K., and Dhami, D. S · 2024
Closest in time.
Learning differentiable logic programs for abstract visual reasoning
Shindo, H., Pfanschilling, V., Dhami, D. S., and Kersting, K · 2024
Closest in time.
Learning by self-explaining
Stammer, W., Friedrich, F., Steinmann, D., Brack, M., Shindo, H., and Kersting, K · 2024
Closest in time.
Towards truly zero-shot compositional visual reasoning with llms as programmers
Stanić, A., Caelles, S., and Tschannen, M · 2024
Closest in time.