Fetching the paper…
Reading the bibliography…
Logical image understanding involves interpreting and reasoning about the relationships and consistency within an image's visual content.
AnomalyGPT: Detecting industrial anomalies using large vision-language models
Gu, Z.; Zhu, B.; Zhu, G.; Chen, Y.; Tang, M.; and Wang, J. 2024a · 1940
Earlier work this paper cites.
Equality and domain closure in first-order databases
Reiter, R. 1980 · 1980
Earlier work this paper cites.
LOF: Identifying Density-based Local Outliers
Breunig, M. M.; Kriegel, H.-P.; Ng, R. T.; and Sander, J. 2000 · 2000
Earlier work this paper cites.
A Mathematical Introduction to Logic
Enderton, H. B. 2001 · 2001
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
MVTec AD — A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection
Bergmann, P.; Fauser, M.; Sattlegger, D.; and Steger, C. 2019 · 2019
Earlier work this paper cites.
Modeling the Distribution of Normal Data in Pre-Trained Deep Features for Anomaly Detection
Rippel, O.; Mertens, P.; and Merhof, D. 2021 · 2020
Earlier work this paper cites.
Emerging Properties in Self-Supervised Vision Transformers
Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; and Joulin, A. 2021 · 2021
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Modeling the Distribution of Normal Data in Pre-Trained Deep Features for Anomaly Detection
Rippel, O.; Mertens, P.; König, E.; and Merhof, D. 2021 · 2021
Earlier work this paper cites.
Beyond Dents and Scratches: Logical Constraints in Unsupervised Anomaly Detection and Localization
Bergmann, P.; Batzner, K.; Fauser, M.; Sattlegger, D.; and Steger, C. 2022 · 2022
Earlier work this paper cites.
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.-W.; Zhu, S.-C.; Tafjord, O.; Clark, P.; and Kalyan, A. 2022 · 2022
Earlier work this paper cites.
Towards Total Recall in Industrial Anomaly Detection
Roth, K.; Pemula, L.; Zepeda, J.; Schölkopf, B.; Brox, T.; and Gehler, P. 2022 · 2022
Earlier work this paper cites.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Cited alongside, same era.
SPot-the-Difference Self-Supervised Pre-training for Anomaly Detection and Segmentation
Zou, Y.; Jeong, J.; Pemula, L.; Zhang, D.; and Dabeer, O. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Chen, X.; Han, Y.; and Zhang, J. 2023 · 2023
Cited alongside, same era.
ImageBind: One Embedding Space To Bind Them All
Girdhar, R.; El-Nouby, A.; Liu, Z.; Singh, M.; Alwala, K. V.; Joulin, A.; and Misra, I. 2023 · 2023
LLaMA: Open and Efficient Foundation Language Models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
Bitnet: Scaling 1-bit transformers for large language models
Wang, H.; Ma, S.; Dong, L.; Huang, S.; Wang, H.; Ma, L.; Yang, F.; Wang, R.; Wu, Y.; and Wei, F. 2023 · 2023
Later among the works it cites.
Visual Chain-of-Thought Prompting for Knowledge-Based Visual Reasoning
Chen, Z.; Zhou, Q.; Shen, Y.; Hong, Y.; Sun, Z.; Gutfreund, D.; and Gan, C. 2024 · 2024
Later among the works it cites.
CVPR, Visual Anomaly and Novelty Detection 2.0 Winner 2024, https://www.hackster.io/contests/openvino2024
Gu, Z.; Zhu, B.; Zhu, G.; Chen, Y.; and Wang, J. 2024b · 2024
Later among the works it cites.
Detecting and Preventing Hallucinations in Large Vision Language Models
Gunjal, A.; Yin, J.; and Bas, E. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
WinCLIP: Zero-/Few-Shot Anomaly Classification and Segmentation
Jeong, J.; Zou, Y.; Kim, T.; Zhang, D.; Ravichandran, A.; and Dabeer, O. 2023 · 2023
Cited alongside, same era.
A Mathematical Investigation of Hallucination and Creativity in GPT Models
Lee, M. 2023 · 2023
Cited alongside, same era.
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Cited alongside, same era.
LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers
Olausson, T. X.; Gu, A.; Lipkin, B.; Zhang, C. E.; Solar-Lezama, A.; Tenenbaum, J. B.; and Levy, R. P. 2023 · 2023
Cited alongside, same era.
Teaching CLIP to Count to Ten
Paiss, R.; Ephrat, A.; Tov, O.; Zada, S.; Mosseri, I.; Irani, M.; and Dekel, T. 2023 · 2023
Cited alongside, same era.
Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning
Pan, L.; Albalak, A.; Wang, X.; and Wang, W. 2023 · 2023
Cited alongside, same era.
Asymmetric Student-Teacher Networks for Industrial Anomaly Detection
Rudolph, M.; Wehrbein, T.; Rosenhahn, B.; and Wandt, B. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Few Shot Part Segmentation Reveals Compositional Logic for Industrial Anomaly Detection
Kim, S.; An, S.; Chikontwe, P.; Kang, M.; Adeli, E.; Pohl, K. M.; and Park, S. H. 2024 · 2024
Later among the works it cites.
Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Shang, Y.; Cai, M.; Xu, B.; Lee, Y. J.; and Yan, Y. 2024 · 2024
Later among the works it cites.
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
Song, Y.; Wang, G.; Li, S.; and Lin, B. Y. 2024 · 2024
Later among the works it cites.
Visionllm: Large language model is also an open-ended decoder for vision-centric tasks
Wang, W.; Chen, Z.; Chen, X.; Wu, J.; Zhu, X.; Zeng, G.; Luo, P.; Lu, T.; Zhou, J.; Qiao, Y.; et al. 2024 · 2024
Later among the works it cites.
SatLM: Satisfiability-Aided Language Models Using Declarative Prompting
Ye, X.; Chen, Q.; Dillig, I.; and Durrett, G. 2024 · 2024
Later among the works it cites.
Good at Captioning, Bad at Counting: Benchmarking GPT-4v on Earth Observation data
Zhang, C.; and Wang, S. 2024 · 2024
Later among the works it cites.
LogiCode: an LLM-Driven Framework for Logical Anomaly Detection
Zhang, Y.; Cao, Y.; Xu, X.; and Shen, W. 2024 · 2024
Later among the works it cites.
AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection
Zhou, Q.; Pang, G.; Tian, Y.; He, S.; and Chen, J. 2024 · 2024
Later among the works it cites.