Fetching the paper…
Reading the bibliography…
State-space models (SSMs) and transformers dominate the language modeling landscape.
Hooker, S · 1911
Earlier work this paper cites.
Representation of events in nerve nets and finite automata
Kleene, S · 1951
Earlier work this paper cites.
On finite monoids having only trivial subgroups
Schützenberger, M · 1965
Earlier work this paper cites.
Bounded-width polynomial-size branching programs recognize exactly those languages in nc1
Barrington, D. A · 1989
Earlier work this paper cites.
Distributed representations, simple recurrent networks, and grammatical structure
Elman, J. L · 1991
Earlier work this paper cites.
On the computational power of neural nets
Siegelmann, H. T. and Sontag, E. D · 1992
Earlier work this paper cites.
Extraction of rules from discrete-time recurrent neural networks
Omlin, C. W. and Giles, C · 1996
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Self-delimiting neural networks
Schmidhuber, J · 2012
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Weston, J., Bordes, A., Chopra, S., Rush, A. M., Van Merriënboer, B., Joulin, A., and Mikolov, T · 2015
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Graves, A · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Deep equilibrium models
Bai, S., Kolter, J. Z., and Koltun, V · 2019
Earlier work this paper cites.
Universal transformers
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, L · 2019
Earlier work this paper cites.
Sequential neural networks as automata
Merrill, W · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
On the Ability and Limitations of Transformers to Recognize Formal Languages
Bhattamishra, S., Ahuja, K., and Goyal, N · 2020
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
A formal hierarchy of rnn architectures
Merrill, W., Weiss, G., Goldberg, Y., Schwartz, R., Smith, N. A., and Yahav, E · 2020
Cited alongside, same era.
Banino, A., Balaguer, J., and Blundell, C · 2021
Cited alongside, same era.
On training implicit models
Geng, Z., Zhang, X.-Y., Bai, S., Wang, Y., and Lin, Z · 2021
Cited alongside, same era.
Implicit representations of meaning in neural language models
Li, B. Z., Nye, M., and Andreas, J · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
Learning associative inference using fast weight memory
Schlag, I., Munkhdalai, T., and Schmidhuber, J · 2021
Simplified state space layers for sequence modeling
Smith, J. T., Warrington, A., and Linderman, S · 2023
Later among the works it cites.
Scaling laws vs model architectures: How does inductive bias influence scaling?
Tay, Y., Dehghani, M., Abnar, S., Chung, H. W., Fedus, W., Rao, J., Narang, S., Tran, V. Q., Yogatama, D., and Metzler, D · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
What algorithms can transformers learn? a study in length generalization
Zhou, H., Bradley, A., Littwin, E., Razin, N., Saremi, O., Susskind, J., Bengio, S., and Nakkiran, P · 2023
Later among the works it cites.
xLSTM: Extended long short-term memory
Beck, M., Pöppel, K., Spanring, M., Auer, A., Prudnikova, O., Kopp, M. K., Klambauer, G., Brandstetter, J., and Hochreiter, S · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Thinking like transformers
Weiss, G., Goldberg, Y., and Yahav, E · 2021
Cited alongside, same era.
Path independent equilibrium models can better exploit test-time computation
Anil, C., Pokle, A., Liang, K., Treutlein, J., Wu, Y., Bai, S., Kolter, J. Z., and Grosse, R. B · 2022
Cited alongside, same era.
Single-sequence protein structure prediction using a language model and deep learning
Chowdhury, R., Bouatta, N., Biswas, S., Floristean, A., Kharkar, A., Roy, R., Rochereau, C., Zhang, J., Church, G. M., Sorger, P. K., and AlQuraishi, M · 2022
Cited alongside, same era.
Learning iterative reasoning through energy minimization
Du, Y., Li, S., Tenenbaum, J., and Mordatch, I · 2022
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Gu, A., Goel, K., and Re, C · 2022
Cited alongside, same era.
Saturated transformers are constant-depth threshold circuits
Merrill, W., Sabharwal, A., and Smith, N. A · 2022
Cited alongside, same era.
Later among the works it cites.
MoEUT: Mixture-of-experts universal transformers
Csordás, R., Irie, K., Schmidhuber, J., Potts, C., and Manning, C. D · 2024
Later among the works it cites.
Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality
Dao, T. and Gu, A · 2024
Later among the works it cites.
A framework for few-shot language model evaluation, 07 2024
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2024
Later among the works it cites.
Can looped transformers learn to implement multi-step gradient descent for in-context learning?
Gatmiry, K., Saunshi, N., Reddi, S. J., Jegelka, S., and Kumar, S · 2024
Later among the works it cites.
Scaling laws and compute-optimal training beyond fixed training durations
Hägele, A., Bakouch, E., Kosson, A., Allal, L. B., Von Werra, L., and Jaggi, M · 2024
Later among the works it cites.
Training large language models to reason in a continuous latent space, 2024
Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J., and Tian, Y · 2024
Later among the works it cites.
Parallelizing non-linear sequential models over the sequence length
Lim, Y. H., Zhu, Q., Selfridge, J., and Kasim, M. F · 2024
Later among the works it cites.
The expressive power of transformers with chain of thought
Merrill, W. and Sabharwal, A · 2024
Later among the works it cites.
The illusion of state in state-space models
Merrill, W., Petty, J., and Sabharwal, A · 2024
Later among the works it cites.
The expressive capacity of state space models: A formal language perspective
Sarrof, Y., Veitsman, Y., and Hahn, M · 2024
Later among the works it cites.
Recurrent transformers trade-off parallelism for length generalization on regular languages
Soulos, P., Terzic, A., Hersche, M., and Rahimi, A · 2024
Later among the works it cites.
What formal languages can transformers express? a survey
Strobl, L., Merrill, W., Weiss, G., Chiang, D., and Angluin, D · 2024
Later among the works it cites.
Evaluating the world model implicit in a generative model
Vafa, K., Chen, J. Y., Rambachan, A., Kleinberg, J., and Mullainathan, S · 2024
Later among the works it cites.
Roadmap on neuromorphic photonics, 2025
Brunner, D., Shastri, B. J., Qadasi, M. A. A., Ballani, H., Barbay, S., Biasi, S., Bienstman, P., Bilodeau, S., Bogaerts, W., Böhm, F., Brennan, G., Buckley, S., Cai, X., Strinati, M. C., Canakci, B., Charbonnier, B., Chemnitz, M., Chen, Y., Cheung, S., Chiles, J., Choi, S., Christodoulides, D. N., Chrostowski, L., Chu, J., Clegg, J. H., Cletheroe, D., Conti, C., Dai, Q., Lauro, L. D., Diamantopoulos, N. P., Dinc, N. U., Ewaniuk, J., Fan, S., Fang, L., Franchi, R., Freire, P., Gentilini, S., Gigan, S., Giorgi, G. L., Gkantsidis, C., Gladrow, J., Goi, E., Goldmann, M., Grabulosa, A., Gu, M., Guo, X., Hejda, M., Horst, F., Hsieh, J. L., Hu, J., Hu, J., Huang, C., Hurtado, A., Jaurigue, L., Kalinin, K. P., Kopae, M. K., Kelly, D. J., Khajavikhan, M., Kremer, H., Laydevant, J., Lederman, J. C., Lee, J., Lenstra, D., Li, G. H. Y., Li, M., Li, Y., Lin, X., Lin, Z., Lis, M., Lüdge, K., Lugnan, A., Lupo, A., Lvovsky, A. I., Manuylovich, E., Marandi, A., Marchesin, F., Massar, S., McCaughan, A. N., McMahon, P. L., Pegios, M. M., Morandotti, R., Moser, C., Moss, D. J., Mukherjee, A., Nikdast, M., Offrein, B. J., Oguz, I., Oripov, B., O’Shea, G., Ozcan, A., Parmigiani, F., Pasricha, S., Pavanello, F., Pavesi, L., Peserico, N., Pickup, L., Pierangeli, D., Pleros, N., Porte, X., Primavera, B. A., Prucnal, P., Psaltis, D., Puts, L., Qiao, F., Rahmani, B., Raineri, F., Ocampo, C. A. R., Robertson, J., Romeira, B., Carmes, C. R., Rotenberg, N., Rowstron, A., Schoenhardt, S., Schwartz, R. L. . T., Shainline, J. M., Shekhar, S., Skalli, A., Sohoni, M. M., Sorger, V. J., Soriano, M. C., Spall, J., Stabile, R., Stiller, B., Sunada, S., Tefas, A., Tossoun, B., Tsakyridis, A., Turitsyn, S. K., der Sande, G. V., Vaerenbergh, T. V., Veraldi, D., Verschaffelt, G., Vlieg, E. A., Wang, H., Wang, T., Wetzstein, G., Wright, L. G., Wu, C., Wu, C., Wu, J., Xia, F., Xu, X., Yang, H., Yao, W., Yildirim, M., Yoo, S. J. B., Youngblood, N., Zambrini, R., Zhang, H., and Zhang, W · 2025
Closest in time.
Unlocking state-tracking in linear RNNs through negative eigenvalues
Grazzi, R., Siems, J., Franke, J. K., Zela, A., Hutter, F., and Pontil, M · 2025
Closest in time.
Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning
Snell, C. V., Lee, J., Xu, K., and Kumar, A · 2025
Closest in time.