Fetching the paper…
Reading the bibliography…
Artificial Intelligence (AI) systems have made remarkable progress, attaining super-human performance across various domains.
Finding and visualizing weaknesses of deep reinforcement learning agents
Rupprecht, C., Ibrahim, C., and Pal, C. J. (2019) · 1904
Earlier work this paper cites.
EDUCE: Explaining model Decisions through Unsupervised Concepts Extraction
Bouchacourt, D. and Denoyer, L. (2019) · 1905
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Tenney, I., Das, D., and Pavlick, E. (2019) · 1905
Earlier work this paper cites.
Explaining classifiers with causal concept effect (CaCE)
Goyal, Y., Feder, A., Shalit, U., and Kim, B. (2019) · 1907
Earlier work this paper cites.
Conservative q-improvement: Reinforcement learning for an interpretable decision-tree policy
Roth, A. M., Topin, N., Jamshidi, P., and Veloso, M. (2019) · 1907
Earlier work this paper cites.
Graying the black box: Understanding dqns
Zahavy, T., Ben-Zrihem, N., and Mannor, S. (2016) · 1908
Earlier work this paper cites.
Counterfactual states for atari agents via generative deep learning
Olson, M. L., Neal, L., Li, F., and Wong, W.-K. (2019) · 1909
Earlier work this paper cites.
Atrey, A., Clary, K., and Jensen, D. (2019) · 1912
Earlier work this paper cites.
Some aspects of nonorthogonal data analysis: Part i. developing prediction equations
Snee, R. D. (1973) · 1973
Earlier work this paper cites.
Comment: Collinearity diagnostics depend on the domain of prediction, the model, and the data
Snee, R. D. and Marquardt, D. W. (1984) · 1984
Earlier work this paper cites.
Fitting equations to data: Computer analysis of multifactor data for scientists and engineers
Robinson, P. J. (1974) · 1986
Earlier work this paper cites.
Case-based evaluation in computer chess
Kerner, Y. (1995) · 1995
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Tibshirani, R. (1996) · 1996
Earlier work this paper cites.
How to Win at Chess
King, D. (2000) · 2000
Earlier work this paper cites.
Distal explanations for explainable reinforcement learning agents
Madumal, P., Miller, T., Sonenberg, L., and Vetere, F. (2020) · 2001
Earlier work this paper cites.
Adversarial tcav–robust and effective interpretation of intermediate layers in neural networks
Soni, R., Shah, N., Seng, C. T., and Moore, J. D. (2020) · 2002
Earlier work this paper cites.
Sreedharan, S., Soni, U., Verma, M., Srivastava, S., and Kambhampati, S. (2020a) · 2002
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N. (2016) · 2003
Earlier work this paper cites.
High-dimensional graphs and variable selection with the lasso
Meinshausen, N. and Bühlmann, P. (2006) · 2006
Earlier work this paper cites.
Debiasing concept bottleneck models with instrumental variables
Bahadori, M. T. and Heckerman, D. E. (2020) · 2007
Earlier work this paper cites.
Koh, P. W., Nguyen, T., Tang, Y. S., Mussmann, S., Pierson, E., Kim, B., and Liang, P. (2020) · 2007
Earlier work this paper cites.
Convex and semi-nonnegative matrix factorizations
Ding, C. H., Li, T., and Jordan, M. I. (2008) · 2008
Earlier work this paper cites.
Assessing game balance with AlphaZero: Exploring alternative rule sets in chess
Tomašev, N., Paquet, U., Hassabis, D., and Kramnik, V. (2020) · 2009
Earlier work this paper cites.
Strategic test suite
Corbit, D., Natarajan, S., and Mosca, F. (2014) · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
CVXPY: A Python-embedded modeling language for convex optimization
Diamond, S. and Boyd, S. (2016) · 2016
Earlier work this paper cites.
Towards deep symbolic reinforcement learning
Garnelo, M., Arulkumaran, K., and Shanahan, M. (2016) · 2016
Earlier work this paper cites.
Why should I trust you?: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C. (2016) · 2016
Earlier work this paper cites.
Visualizing dynamics: from t-sne to semi-mdps
Zrihem, N. B., Zahavy, T., and Mannor, S. (2016) · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A. (2017) · 2017
Earlier work this paper cites.
Understanding Chess with Explainable AI
DecodeChess (2017) · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P. (2017) · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S.-I. (2017) · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2017) · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q. (2017) · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B. (2018) · 2018
Earlier work this paper cites.
A rewriting system for convex optimization problems
Agrawal, A., Verschueren, R., Diamond, S., and Boyd, S. (2018) · 2018
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
Alvarez-Melis, D. and Jaakkola, T. S. (2018) · 2018
Earlier work this paper cites.
Verifiable reinforcement learning via policy extraction
Bastani, O., Pu, Y., and Solar-Lezama, A. (2018) · 2018
Earlier work this paper cites.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
Conneau, A., Kruszewski, G., Lample, G., Barrault, L., and Baroni, M. (2018) · 2018
Earlier work this paper cites.
Towards symbolic reinforcement learning with common sense
d’Avila Garcez, A., Dutra, A. R. R., and Alonso, E. (2018) · 2018
Earlier work this paper cites.
Open-ended learning: A conceptual framework based on representational redescription
Doncieux, S., Filliat, D., Díaz-Rodríguez, N., Hospedales, T., Duro, R., Coninx, A., Roijers, D. M., Girard, B., Perrin, N., and Sigaud, O. (2018) · 2018
Earlier work this paper cites.
Unsupervised video object segmentation for deep reinforcement learning
Goel, V., Weng, J., and Poupart, P. (2018) · 2018
Cited alongside, same era.
Regression concept vectors for bidirectional explanations in histopathology
Graziani, M., Andrearczyk, V., and Müller, H. (2018) · 2018
Cited alongside, same era.
Visualizing and understanding atari agents
Greydanus, S., Koul, A., Dodge, J., and Fern, A. (2018) · 2018
Cited alongside, same era.
Learning to generate move-by-move commentary for chess games from large-scale social forum data
Jhamtani, H., Gangal, V., Hovy, E., Neubig, G., and Berg-Kirkpatrick, T. (2018) · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al. (2018) · 2018
Cited alongside, same era.
Learning finite state representations of recurrent policy networks
Concept-based model explanations for electronic health records
Mincu, D., Loreaux, E., Hou, S., Baur, S., Protsyuk, I., Seneviratne, M., Mottram, A., Tomasev, N., Karthikesalingam, A., and Schrouff, J. (2021) · 2021
Later among the works it cites.
Iterative bounding mdps: Learning interpretable policies via non-interpretable methods
Topin, N., Milani, S., Fang, F., and Veloso, M. (2021) · 2021
Later among the works it cites.
From ”where” to ”what”: Towards human-understandable explanations through concept relevance propagation
Achtibat, R., Dreyer, M., Eisenbraun, I., Bosse, S., Wiegand, T., Samek, W., and Lapuschkin, S. (2022) · 2022
Later among the works it cites.
Concept gradient: Concept-based interpretation without linear assumption
Bai, A., Yeh, C.-K., Ravikumar, P., Lin, N. Y. C., and Hsieh, C.-J. (2022) · 2022
Later among the works it cites.
Impossibility theorems for feature attribution
Bilodeau, B., Jaques, N., Koh, P. W., and Kim, B. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Koul, A., Greydanus, S., and Fern, A. (2018) · 2018
Cited alongside, same era.
Object-sensitive deep reinforcement learning
Li, Y., Sycara, K., and Iyer, R. (2018) · 2018
Cited alongside, same era.
Toward interpretable deep reinforcement learning with linear model u-trees
Liu, G., Schulte, O., Zhu, W., and Li, Q. (2019) · 2018
Cited alongside, same era.
A theoretical explanation for perplexing behaviors of backpropagation-based visualizations
Nie, W., Zhang, Y., and Patel, A. (2018) · 2018
Cited alongside, same era.
S-rl toolbox: Environments, datasets and evaluation metrics for state representation learning
Raffin, A., Hill, A., Traoré, R., Lesort, T., Díaz-Rodríguez, N., and Filliat, D. (2018) · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D. (2018) · 2018
Cited alongside, same era.
Relational deep reinforcement learning
Zambaldi, V., Raposo, D., Santoro, A., Bapst, V., Li, Y., Babuschkin, I., Tuyls, K., Reichert, D., Lillicrap, T., Lockhart, E., Shanahan, M., Langston, V., Pascanu, R., Botvinick, M., Vinyals, O., and Battaglia, P. (2018) · 2018
Cited alongside, same era.
Impossibility theorems for feature attribution
Bilodeau, B., Jaques, N., Koh, P. W., and Kim, B. (2022) · 2022
Later among the works it cites.
Concept activation regions: A generalized framework for concept-based explanations
Crabbé, J. and van der Schaar, M. (2022) · 2022
Later among the works it cites.
Explainable reinforcement learning via model transforms
Finkelstein, M., Liu, L., Schlot, N. L., Kolumbus, Y., Parkes, D. C., Rosenshein, J. S., and Keren, S. (2022) · 2022
Later among the works it cites.
Where, when & which concepts does AlphaZero learn? Lessons from the game of Hex
Forde, J. Z., Lovering, C., Konidaris, G., Pavlick, E., and Littman, M. L. (2022) · 2022
Later among the works it cites.
Reccover: Detecting causal confusion for explainable reinforcement learning
Gajcin, J. and Dusparic, I. (2022) · 2022
Later among the works it cites.
Alphazero ideas
González-Díaz, J. and Palacios-Huerta, I. (2022) · 2022
Later among the works it cites.
The role of explainability in assuring safety of machine learning in healthcare
Jia, Y., McDermid, J., Lawton, T., and Habli, I. (2022) · 2022
Later among the works it cites.
Beyond interpretability: developing a language to shape our relationships with ai
Kim, B. (2022) · 2022
Later among the works it cites.
Explainability in reinforcement learning: perspective and position
Krajna, A., Brcic, M., Lipic, T., and Doncevic, J. (2022) · 2022
Later among the works it cites.
Are alphazero-like agents robust to adversarial perturbations?
Lan, L.-C., Zhang, H., Wu, T.-R., Tsai, M.-Y., Wu, I.-C., and Hsieh, C.-J. (2022) · 2022
Later among the works it cites.
Acquisition of chess knowledge in alphazero
McGrath, T., Kapishnikov, A., Tomašev, N., Pearce, A., Wattenberg, M., Hassabis, D., Kim, B., Paquet, U., and Kramnik, V. (2022) · 2022
Later among the works it cites.
A survey of explainable reinforcement learning
Milani, S., Topin, N., Veloso, M., and Fang, F. (2022) · 2022
Later among the works it cites.
Beyond rewards: a hierarchical perspective on offline multiagent behavioral analysis
Omidshafiei, S., Kapishnikov, A., Assogba, Y., Dixon, L., and Kim, B. (2022) · 2022
Later among the works it cites.
Reimagining chess with AlphaZero
Tomašev, N., Paquet, U., Hassabis, D., and Kramnik, V. (2022) · 2022
Later among the works it cites.
Understanding game-playing agents with natural language annotations
Tomlin, N., He, A., and Klein, D. (2022) · 2022
Later among the works it cites.
Explainable deep reinforcement learning: state of the art and challenges
Vouros, G. A. (2022) · 2022
Later among the works it cites.
Interview with gm chuchelov - caruana’s coach
Alisa Melekhina (2014) · 2023
Closest in time.
Syzygy tablebase
Bojun Guo (2023) · 2023
Closest in time.
State2explanation: Concept-based explanations to benefit agent learning and user understanding
Das, D., Chernova, S., and Kim, B. (2023) · 2023
Closest in time.
Explainable reinforcement learning for broad-xai: a conceptual framework and survey
Dazeley, R., Vamplew, P., and Cruz, F. (2023) · 2023
Closest in time.
Counterfactual explanation policies in rl
Deshmukh, S. V., R, S., Vijay, S., Subramanian, J., and Agarwal, C. (2023) · 2023
Closest in time.
Chessgpt: Bridging policy learning and language modeling
Feng, X., Luo, Y., Wang, Z., Tang, H., Yang, M., Shao, K., Mguni, D., Du, Y., and Wang, J. (2023) · 2023
Closest in time.
Fide handbook c.02
FIDE (2019) · 2023
Closest in time.
python-chess
Fiekas, N. (2023) · 2023
Closest in time.
Finding neurons in a haystack: Case studies with sparse probing
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D. (2023) · 2023
Closest in time.
Deep explainable relational reinforcement learning: A neuro-symbolic approach
Hazra, R. and De Raedt, L. (2023) · 2023
Closest in time.
Encyclopedia of chess openings
LiChess (2023) · 2023
Closest in time.
Actually, othello-gpt has a linear emergent world model
Nanda, N. (2023) · 2023
Closest in time.
Unveiling concepts learned by a world-class chess-playing agent
Pálsson, A. and Björnsson, Y. (2023) · 2023
Closest in time.
Overlooked factors in concept-based explanations: Dataset choice, concept learnability, and human capability
Ramaswamy, V. V., Kim, S. S. Y., Fong, R., and Russakovsky, O. (2023) · 2023
Closest in time.
Stockfish Chess
Stockfish Community (2018) · 2023
Closest in time.
Chess rating system
Wikipedia contributors (2023a) · 2023
Closest in time.
Chess strategy
Wikipedia contributors (2023b) · 2023
Closest in time.
Chess tactic
Wikipedia contributors (2023c) · 2023
Closest in time.
Causal proxy models for concept-based model explanations
Wu, Z., D’Oosterlinck, K., Geiger, A., Zur, A., and Potts, C. (2023) · 2023
Closest in time.
Leveraging reward consistency for interpretable feature discovery in reinforcement learning
Yang, Q., Wang, H., Tong, M., Shi, W., Huang, G., and Song, S. (2023) · 2023
Closest in time.
Diversifying ai: Towards creative chess with alphazero
Zahavy, T., Veeriah, V., Hou, S., Waugh, K., Lai, M., Leurent, E., Tomasev, N., Schut, L., Hassabis, D., and Singh, S. (2023) · 2023
Closest in time.