Fetching the paper…
Reading the bibliography…
Many complex problems are naturally understood in terms of symbolic concepts.
Explaining classifiers with causal concept effect
Yash Goyal, Amir Feder, Uri Shalit, and Been Kim. 2019a · 1907
Earlier work this paper cites.
Psychologism and behaviorism
Ned Block. 1981 · 1981
Earlier work this paper cites.
Vision: A computational investigation into the human representation and processing of visual information
David Marr. 1982 · 1982
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Jerry A Fodor and Zenon W Pylyshyn. 1988 · 1988
Earlier work this paper cites.
Concepts: Where cognitive science went wrong
Jerry A Fodor. 1998 · 1998
Earlier work this paper cites.
Concepts: core readings
Eric Margolis, Stephen Laurence, et al. 1999 · 1999
Earlier work this paper cites.
Causal models: How people think about the world and its alternatives
Steven Sloman. 2005 · 2005
Earlier work this paper cites.
When bert forgets how to pos: Amnesic probing of linguistic properties and mlm predictions
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2020 · 2006
Earlier work this paper cites.
Core knowledge
Elizabeth S Spelke and Katherine D Kinzler. 2007 · 2007
Earlier work this paper cites.
The Origin of Concepts
Susan Carey. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020 · 2010
Earlier work this paper cites.
Alex Warstadt, Yian Zhang, Haau-Sing Li, Haokun Liu, and Samuel R Bowman. 2020 · 2010
Earlier work this paper cites.
Measuring and reducing gendered correlations in pre-trained models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020 · 2010
Earlier work this paper cites.
Debugging tests for model explanations
Julius Adebayo, Michael Muelly, Ilaria Liccardi, and Been Kim. 2020 · 2011
Earlier work this paper cites.
Vector space models of word meaning and phrase meaning: A survey
Katrin Erk. 2012 · 2012
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Quoc V Le. 2013 · 2013
Earlier work this paper cites.
Overfeat: Integrated recognition, localization and detection using convolutional networks
Pierre Sermanet, David Eigen, Xiang Zhang, Michaël Mathieu, Rob Fergus, and Yann LeCun. 2013 · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2014 · 2015
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016 · 2016
Cited alongside, same era.
Probing for semantic evidence of composition by means of simple classification tasks
Allyson Ettinger, Ahmed Elgohary, and Philip Resnik. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
Diagnostic Classifiers: Revealing how Neural Networks Process Hierarchical Structure
Sara Veldhoen, Dieuwke Hupkes, and Willem Zuidema. 2016 · 2016
Cited alongside, same era.
Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017 · 2017
Can neural networks understand monotonicity reasoning?
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019 · 2019
Later among the works it cites.
Distributional semantics and linguistic theory
Gemma Boleda. 2020 · 2020
Later among the works it cites.
Probing linguistic systematicity
Emily Goodwin, Koustuv Sinha, and Timothy J. O’Donnell. 2020 · 2020
Later among the works it cites.
Reducing sentiment bias in language models via counterfactual evaluation
Po-Sen Huang, Huan Zhang, Ray Jiang, Robert Stanforth, Johannes Welbl, Jack Rae, Vishal Maini, Dani Yogatama, and Pushmeet Kohli. 2020 · 2020
Later among the works it cites.
Meaning Holism
Henry Jackman. 2020 · 2020
Later among the works it cites.
COGS: A compositional generalization challenge based on semantic interpretation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Interpretable explanations of black boxes by meaningful perturbation
Ruth C Fong and Andrea Vedaldi. 2017 · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017 · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Explaining image classifiers by counterfactual generation
Chun-Hao Chang, Elliot Creager, Anna Goldenberg, and David Duvenaud. 2018 · 2018
Cited alongside, same era.
Visualisation and ‘diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. 2018 · 2018
Cited alongside, same era.
Najoung Kim and Tal Linzen. 2020 · 2020
Later among the works it cites.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020 · 2020
Later among the works it cites.
Syntactic structure from deep learning
Tal Linzen and Marco Baroni. 2021 · 2021
Later among the works it cites.
Predicting inductive biases of pre-trained models
Charles Lovering, Rohan Jha, Tal Linzen, and Ellie Pavlick. 2021 · 2021
Later among the works it cites.
William Merrill, Yoav Goldberg, Roy Schwartz, and Noah A Smith. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Later among the works it cites.
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
What if this modified that? syntactic interventions with counterfactual embeddings
Mycal Tucker, Peng Qian, and Roger Levy. 2021 · 2021
Later among the works it cites.
Locating and editing factual knowledge in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Closest in time.
Aaron Mueller, Robert Frank, Tal Linzen, Luheng Wang, and Sebastian Schuster. 2022 · 2022
Closest in time.
Semantic structure in deep learning
Ellie Pavlick. 2022 · 2022
Closest in time.