Fetching the paper…
Reading the bibliography…
Compositional generalization is crucial for artificial intelligence agents to solve complex vision-language reasoning tasks.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Köpf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; and Chintala, S. 2019 · 1912
Earlier work this paper cites.
Prompting is programming: A query language for large language models
Beurer-Kellner, L.; Fischer, M.; and Vechev, M. 2023 · 1969
Earlier work this paper cites.
Compositionality
Partee, B.; et al. 1984 · 1984
Earlier work this paper cites.
Pearson Correlation Coefficient , 1–4
Benesty, J.; Chen, J.; Huang, Y.; and Cohen, I. 2009 · 2009
Earlier work this paper cites.
Glove: Global Vectors for Word Representation
Pennington, J.; Socher, R.; and Manning, C. 2014 · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. B. 2017 · 2017
Earlier work this paper cites.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Honnibal, M.; and Montani, I. 2017 · 2017
Earlier work this paper cites.
Acquisition of Localization Confidence for Accurate Object Detection
Jiang, B.; Luo, R.; Mao, J.; Xiao, T.; and Jiang, Y. 2018 · 2018
Earlier work this paper cites.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Yi, K.; Wu, J.; Gan, C.; Torralba, A.; Kohli, P.; and Tenenbaum, J. 2018 · 2018
Earlier work this paper cites.
The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
Mao, J.; Gan, C.; Kohli, P.; Tenenbaum, J. B.; and Wu, J. 2019 · 2019
Earlier work this paper cites.
Systematic Generalization on gSCAN with Language Conditioned Embedding
Gao, T.; Huang, Q.; and Mooney, R. 2020 · 2020
Earlier work this paper cites.
Inference-Masked Loss for Deep Structured Output Learning
Guo, Q.; Faghihi, H. R.; Zhang, Y.; Uszok, A.; and Kordjamshidi, P. 2020 · 2020
Earlier work this paper cites.
Compositionality Decomposed: How do Neural Networks Generalise? (Extended Abstract)
Hupkes, D.; Dankers, V.; Mul, M.; and Bruni, E. 2020 · 2020
Earlier work this paper cites.
A benchmark for systematic generalization in grounded language understanding
Ruis, L.; Andreas, J.; Baroni, M.; Bouchacourt, D.; and Lake, B. 2020 · 2020
Earlier work this paper cites.
The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of Transformers
Csordás, R.; Irie, K.; and Schmidhuber, J. 2021 · 2021
Cited alongside, same era.
DomiKnowS: A Library for Integration of Symbolic Domain Knowledge in Deep Learning
Faghihi, H. R.; Guo, Q.; Uszok, A.; Nafar, A.; Raisi, E.; and Kordjamshidi, P. 2021 · 2021
Cited alongside, same era.
Inducing Transformer’s Compositional Generalization Ability via Auxiliary Sequence Prediction Tasks
Jiang, Y.; and Bansal, M. 2021 · 2021
Cited alongside, same era.
MDETR - Modulated Detection for End-to-End Multi-Modal Understanding
Kamath, A.; Singh, M.; LeCun, Y.; Synnaeve, G.; Misra, I.; and Carion, N. 2021 · 2021
Cited alongside, same era.
Compositional Networks Enable Systematic Generalization for Grounded Language Understanding
Kuo, Y.-L.; Katz, B.; and Barbu, A. 2021 · 2021
Cited alongside, same era.
GLUECons: A Generic Benchmark for Learning under Constraints
Rajaby Faghihi, H.; Nafar, A.; Zheng, C.; Mirzaee, R.; Zhang, Y.; Uszok, A.; Wan, A.; Premsri, T.; Roth, D.; and Kordjamshidi, P. 2023 · 2023
Later among the works it cites.
ViperGPT: Visual Inference via Python Execution for Reasoning
Surís, D.; Menon, S.; Vondrick, C.; and . 2023 · 2023
Later among the works it cites.
Florence-2: Advancing a unified representation for a variety of vision tasks
Xiao, B.; Wu, H.; Xu, W.; Dai, X.; Hu, H.; Lu, Y.; Zeng, M.; Liu, C.; and Yuan, L. 2023 · 2023
Later among the works it cites.
MetaReVision: Meta-Learning with Retrieval for Visually Grounded Compositional Concept Acquisition
Xu, G.; Kordjamshidi, P.; and Chai, J. 2023 · 2023
Later among the works it cites.
Meta Compositional Referring Expression Segmentation
Xu, L.; Huang, M. H.; Shang, X.; Yuan, Z.; Sun, Y.; and Liu, J. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ontañón, S.; Ainslie, J.; Cvicek, V.; and Fisher, Z. 2021 · 2021
Cited alongside, same era.
Systematic Generalization on gSCAN: What is Nearly Solved and What is Next?
Qiu, L.; Hu, H.; Zhang, B.; Shaw, P.; and Sha, F. 2021 · 2021
Cited alongside, same era.
ReaSCAN: Compositional Reasoning in Language Grounding
Wu, Z.; Kreiss, E.; Ong, D. C.; and Potts, C. 2021 · 2021
Cited alongside, same era.
When Can Transformers Ground and Compose: Insights from Compositional Generalization Benchmarks
Sikarwar, A.; Patel, A.; and Goyal, N. 2022 · 2022
Cited alongside, same era.
Generalization Differences between End-to-End and Neuro-Symbolic Vision-Language Reasoning Systems
Zhu, W.; Thomason, J.; and Jia, R. 2022 · 2022
Cited alongside, same era.
Openflamingo: An open-source framework for training large autoregressive vision-language models
Awadalla, A.; Gao, I.; Gardner, J.; Hessel, J.; Hanafy, Y.; Zhu, W.; Marathe, K.; Bitton, Y.; Gadre, S.; Sagawa, S.; et al. 2023 · 2023
Cited alongside, same era.
Binding Language Models in Symbolic Languages
Cheng, Z.; Xie, T.; Shi, P.; Li, C.; Nadkarni, R.; Hu, Y.; Xiong, C.; Radev, D.; Ostendorf, M.; Zettlemoyer, L.; Smith, N. A.; and Yu, T. 2023 · 2023
Cited alongside, same era.
Do Vision-Language Pretrained Models Learn Composable Primitive Concepts?
Yun, T.; Bhalla, U.; Pavlick, E.; and Sun, C. 2023 · 2023
Later among the works it cites.
Parsel: Algorithmic Reasoning with Language Models by Composing Decompositions
Zelikman, E.; Huang, Q.; Poesia, G.; Goodman, N.; and Haber, N. 2023 · 2023
Later among the works it cites.
VLN-Trans: Translator for the Vision and Language Navigation Agent
Zhang, Y.; and Kordjamshidi, P. 2023 · 2023
Later among the works it cites.
Dubey, A.; Jauhri, A.; and Others. 2024 · 2024
Closest in time.
Prompt2DeModel: Declarative Neuro-Symbolic Modeling with Natural Language
Faghihi, H. R.; Nafar, A.; Uszok, A.; Karimian, H.; and Kordjamshidi, P. 2024 · 2024
Closest in time.
What’s left? concept grounding with logic-enhanced foundation models
Hsu, J.; Mao, J.; Tenenbaum, J.; and Wu, J. 2024 · 2024
Closest in time.
LLaVA-NeXT: Improved reasoning, OCR, and world knowledge
Liu, H.; Li, C.; Li, Y.; Li, B.; Zhang, Y.; Shen, S.; and Lee, Y. J. 2024 · 2024
Closest in time.
A Survey on Compositional Learning of AI Models: Theoretical and Experimental Practices
Sinha, S.; Premsri, T.; and Kordjamshidi, P. 2024 · 2024
Closest in time.
Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution
Wang, P.; Bai, S.; Tan, S.; Wang, S.; Fan, Z.; Bai, J.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Fan, Y.; Dang, K.; Du, M.; Ren, X.; Men, R.; Liu, D.; Zhou, C.; Zhou, J.; and Lin, J. 2024 · 2024
Closest in time.
GIPCOL: Graph-Injected Soft Prompting for Compositional Zero-Shot Learning
Xu, G.; Kordjamshidi, P.; and Chai, J. 2024 · 2024
Closest in time.
Vision-and-language navigation today and tomorrow: A survey in the era of foundation models
Zhang, Y.; Ma, Z.; Li, J.; Qiao, Y.; Wang, Z.; Chai, J.; Wu, Q.; Bansal, M.; and Kordjamshidi, P. 2024 · 2024
Closest in time.