Fetching the paper…
Reading the bibliography…
Large language models (LLMs) excel at generating fluent text, but their internal reasoning remains opaque and difficult to control.
Pattern recognition as rule-guided inductive inference
Ryszard S. Michalski · 1980
Earlier work this paper cites.
Foundations of Logic Programming
John W. Lloyd · 1984
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach , chapter 9.3 Forward Chaining, pp. 330–337
Stuart Russell and Peter Norvig · 1995
Earlier work this paper cites.
Machine learning, international edition
Tom Michael Mitchell · 1997
Earlier work this paper cites.
The aleph manual
Ashwin Srinivasan · 2001
Earlier work this paper cites.
Structure learning of probabilistic logic programs by searching the clause space
Elena Bellodi and Fabrizio Riguzzi · 2013
Earlier work this paper cites.
Inducing probabilistic relational rules from probabilistic examples
Luc De Raedt, Anton Dries, Ingo Thon, Guy Van den Broeck, Mathias Verbeke, Q Yang, and M Wooldridge · 2015
Earlier work this paper cites.
Neural-symbolic learning and reasoning: contributions and challenges
Artur d’Avila Garcez, Tarek R Besold, Luc De Raedt, Peter Földiak, Pascal Hitzler, Thomas Icard, Kai-Uwe Kühnberger, Luis C Lamb, Risto Miikkulainen, and Daniel L Silver · 2015
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory Sayres · 2018
Earlier work this paper cites.
Deepproblog: Neural probabilistic logic programming
Robin Manhaeve, Sebastijan Dumancic, Angelika Kimmig, Thomas Demeester, and Luc De Raedt · 2018
Earlier work this paper cites.
Towards automatic concept-based explanations
Amirata Ghorbani, James Wexler, James Y Zou, and Been Kim · 2019
Earlier work this paper cites.
Artificial intelligence, values and alignment
Iason Gabriel · 2020
Earlier work this paper cites.
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Earlier work this paper cites.
The next decade in ai: four steps towards robust artificial intelligence
Gary Marcus · 2020
Earlier work this paper cites.
Language models are few-shot learners
OpenAI · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, et al · 2021
Earlier work this paper cites.
Learning programs by learning from failures
Andrew Cropper and Rolf Morel · 2021
Earlier work this paper cites.
Neuro-symbolic artificial intelligence
Md Kamruzzaman Sarker, Lu Zhou, Aaron Eberhart, and Pascal Hitzler · 2021
Earlier work this paper cites.
Slash: embracing probabilistic circuits into neural answer set programming
Arseny Skryagin, Wolfgang Stammer, Daniel Ochs, Devendra Singh Dhami, and Kristian Kersting · 2021
Earlier work this paper cites.
Right for the right concept: Revising neuro-symbolic concepts by interacting with their explanations
Wolfgang Stammer, Patrick Schramowski, and Kristian Kersting · 2021
Earlier work this paper cites.
Selection-inference: Exploiting large language models for interpretable logical reasoning
Antonia Creswell, Murray Shanahan, and Irina Higgins · 2022
Earlier work this paper cites.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Earlier work this paper cites.
Addressing leakage in concept bottleneck models
Marton Havasi, Sonali Parbhoo, and Finale Doshi-Velez · 2022
Earlier work this paper cites.
Symbols as a lingua franca for bridging human-ai chasm for explainable and advisable AI systems
Subbarao Kambhampati, Sarath Sreedharan, Mudit Verma, Yantian Zha, and Lin Guan · 2022
Earlier work this paper cites.
The third AI summer: AAAI Robert S. Engelmore memorial lecture
Henry Kautz · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei et al · 2022
Earlier work this paper cites.
Interpretable neural-symbolic concept reasoning
Pietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Mateo Espinosa Zarlenga, Lucie Charlotte Magister, Alberto Tonda, Pietro Lió, Frederic Precioso, Mateja Jamnik, and Giuseppe Marra · 2023
Cited alongside, same era.
SEGA: Instructing text-to-image models using semantic guidance
Manuel Brack, Felix Friedrich, Dominik Hintersdorf, Lukas Struppek, Patrick Schramowski, and Kristian Kersting · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Cited alongside, same era.
Revision transformers: Instructing language models to change their values
Felix Friedrich, Wolfgang Stammer, Patrick Schramowski, and Kristian Kersting · 2023
Cited alongside, same era.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Learning differentiable logic programs for abstract visual reasoning
Hikaru Shindo, Viktor Pfanschilling, Devendra Singh Dhami, and Kristian Kersting · 2024
Later among the works it cites.
Neural concept binder
Wolfgang Stammer, Antonia Wüst, David Steinmann, and Kristian Kersting · 2024
Later among the works it cites.
Learning to intervene on concept bottlenecks
David Steinmann, Wolfgang Stammer, Felix Friedrich, and Kristian Kersting · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Gemma Team · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Chi Zhang, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Cited alongside, same era.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng · 2023
Cited alongside, same era.
Label-free concept bottleneck models
Tuomas Oikarinen, Subhro Das, Lam M. Nguyen, and Tsui-Wei Weng · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Cited alongside, same era.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Abulhair Saparov and He He · 2023
Cited alongside, same era.
Testing the general deductive reasoning capacity of large language models using OOD examples
Abulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi, Seyed Mehran Kazemi, Najoung Kim, and He He · 2023
Cited alongside, same era.
α \alpha ilp: thinking visual scenes as differentiable logic programs
Hikaru Shindo, Viktor Pfanschilling, Devendra Singh Dhami, and Kristian Kersting · 2023
Cited alongside, same era.
Pix2code: Learning to compose neural visual concepts as programs
Antonia Wüst, Wolfgang Stammer, Quentin Delfosse, Devendra Singh Dhami, and Kristian Kersting · 2024
Later among the works it cites.
On memorization of large language models in logical reasoning
Chulin Xie, Yangsibo Huang, Chiyuan Zhang, Da Yu, Xinyun Chen, Bill Yuchen Lin, Bo Li, Badih Ghazi, and Ravi Kumar · 2024
Later among the works it cites.
A survey on neural-symbolic learning systems
Dongran Yu, Bo Yang, Dayou Liu, Hui Wang, and Shirui Pan · 2024
Later among the works it cites.
Learning multi-level features with matryoshka sparse autoencoders
Bart Bussmann, Noa Nabeshima, Adam Karvonen, and Neel Nanda · 2025
Closest in time.
Saemnesia: Erasing concepts in diffusion models with sparse autoencoders
Enrico Cassano, Riccardo Renzulli, Marco Nurisso, Mirko Zaffaroni, Alan Perotti, and Marco Grangetto · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
Deep reinforcement learning agents are not even close to human intelligence
Quentin Delfosse, Jannis Blüml, Fabian Tatai, Théo Vincent, Bjarne Gregori, Elisabeth Dillies, Jan Peters, Constantin A. Rothkopf, and Kristian Kersting · 2025
Closest in time.
Archetypal SAE: Adaptive and stable dictionary learning for concept extraction in large vision models
Thomas Fel, Ekdeep Singh Lubana, Jacob S. Prince, Matthew Kowal, Victor Boutin, Isabel Papadimitriou, Binxu Wang, Martin Wattenberg, Demba E. Ba, and Talia Konkle · 2025
Closest in time.
Measuring and guiding monosemanticity
Ruben Härle, Felix Friedrich, Manuel Brack, Björn Deiseroth, Stephan Waeldchen, Patrick Schramowski, and Kristian Kersting · 2025
Closest in time.
Slr: Automated synthesis for scalable logical reasoning
Lukas Helff, Ahmad Omar, Felix Friedrich, Antonia Wüst, Hikaru Shindo, Rupert Mitchell, Tim Woydt, Patrick Schramowski, Wolfgang Stammer, and Kristian Kersting · 2025
Closest in time.
Saebench: A comprehensive benchmark for sparse autoencoders in language model interpretability
Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Isaac Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Callum Stuart McDougall, Kola Ayonrinde, et al · 2025
Closest in time.
Sparse autoencoders do not find canonical units of analysis
Patrick Leask, Bart Bussmann, Michael T Pearce, Joseph Isaac Bloom, Curt Tigges, Noura Al Moubayed, Lee Sharkey, and Neel Nanda · 2025
Closest in time.
Sparse autoencoders learn monosemantic features in vision-language models
Mateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge Belongie, and Zeynep Akata · 2025
Closest in time.
Use sparse autoencoders to discover unknown concepts, not to act on known concepts
Kenny Peng, Rajiv Movva, Jon M. Kleinberg, Emma Pierson, and Nikhil Garg · 2025
Closest in time.
Large language models meet symbolic provers for logical reasoning evaluation
Chengwen Qi, Ren Ma, Bowen Li, He Du, Binyuan Hui, Jinwang Wu, Yuanjun Laili, and Conghui He · 2025
Closest in time.
The birth of knowledge: Emergent features across time, space, and scale in large language models
Shashata Sawmya, Micah Adler, and Nir Shavit · 2025
Closest in time.
Bridging the human–ai knowledge gap through concept discovery and transfer in alphazero
Lisa Schut, Nenad Tomašev, Thomas McGrath, Demis Hassabis, Ulrich Paquet, and Been Kim · 2025
Closest in time.
Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh, Maxwell Horton, Samy Bengio, and Mehrdad Farajtabar · 2025
Closest in time.
The value of symbolic concepts for ai explanations and interactions
Wolfgang Stammer · 2025
Closest in time.
Fodor and pylyshyn’s legacy - still no human-like systematic compositionality in neural networks
Tim Woydt, Moritz Willig, Antonia Wüst, Lukas Helff, Wolfgang Stammer, Constantin A. Rothkopf, and Kristian Kersting · 2025
Closest in time.
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al · 2025
Closest in time.