Fetching the paper…
Reading the bibliography…
One way to interpret the behavior of a blackbox recurrent neural network (RNN) is to extract from it a more interpretable discrete computational model, like a finite state machine, that captures its behavior.
Three models for the description of language
Chomsky, N · 1956
Earlier work this paper cites.
Automata studies
Kleene, S. C., Shannon, C. E., and McCarthy, J · 1956
Earlier work this paper cites.
Some universal elements for finite automata
Minsky, M. L · 1956
Earlier work this paper cites.
Counter machines and counter languages
Fischer, P. C., Meyer, A. R., and Rosenberg, A. L · 1968
Earlier work this paper cites.
An n log n n\log n algorithm for minimizing states in a finite automaton
Hopcroft, J · 1971
Earlier work this paper cites.
Complexity of automaton identification from given data
Gold, E. M · 1978
Earlier work this paper cites.
Dynamic construction of finite-state automata from examples using hill-climbing
Tomita, M · 1982
Earlier work this paper cites.
Learning regular sets from queries and counterexamples
Angluin, D · 1987
Earlier work this paper cites.
Finding structure in time
Elman, J. L · 1990
Cited alongside, same era.
Inferring Regular Languages in Polynomial Time , pp. 49–61
Oncina, J. and García, P · 1992
Cited alongside, same era.
On the computational power of neural nets
Siegelmann, H. T. and Sontag, E · 1992
Cited alongside, same era.
Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J · 1997
Cited alongside, same era.
Results of the abbadingo one dfa learning competition and a new evidence-driven state merging algorithm
Lang, K. J., Pearlmutter, B. A., and Price, R. A · 1998
Cited alongside, same era.
On state merging in grammatical inference: A statistical approach for dealing with noisy data
Sebban, M. and Janodet, J.-C · 2003
Cited alongside, same era.
On the information bottleneck theory of deep learning
Saxe, A. M., Bansal, Y., Dapello, J., Advani, M., Kolchinsky, A., Tracey, B. D., and Cox, D. D · 2018
Later among the works it cites.
An empirical evaluation of rule extraction from recurrent neural networks, 2018
Wang, Q., Zhang, K., au2, A. G. O. I., Xing, X., Liu, X., and Giles, C. L · 2018
Later among the works it cites.
Sequential neural networks as automata
Merrill, W · 2019
Later among the works it cites.
A formal hierarchy of RNN architectures
Merrill, W., Weiss, G., Goldberg, Y., Schwartz, R., Smith, N. A., and Yahav, E · 2020
Later among the works it cites.
Extracting weighted automata for approximate minimization in language modelling
Lacroce, C., Panangaden, P., and Rabusseau, G · 2021
Later among the works it cites.
Effects of parameter norm growth during transformer training: Inductive bias from gradient descent
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the properties of neural machine translation: Encoder-decoder approaches, 2014
Cho, K., van Merrienboer, B., Bahdanau, D., and Bengio, Y · 2014
Cited alongside, same era.
Opening the black box of deep neural networks via information
Shwartz-Ziv, R. and Tishby, N · 2017
Cited alongside, same era.
Visualizing and understanding recurrent networks, 2015a
Karpathy, A., Johnson, J., and Fei-Fei, L
Cited in the paper.
Visualizing and understanding recurrent networks, 2015b
Karpathy, A., Johnson, J., and Fei-Fei, L
Cited in the paper.
On the practical computational power of finite precision RNNs for language recognition, 2018a
Weiss, G., Goldberg, Y., and Yahav, E
Cited in the paper.
Extracting automata from recurrent neural networks using queries and counterexamples
Weiss, G., Goldberg, Y., and Yahav, E
Cited in the paper.
Merrill, W., Ramanujan, V., Goldberg, Y., Schwartz, R., and Smith, N. A · 2021
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2021
Later among the works it cites.