Fetching the paper…
Reading the bibliography…
The need to develop a general framework for architecture analysis is becoming increasingly important, given the expanding design space of sequence models.
Dupont, E., Doucet, A., and Teh, Y. W · 1904
Earlier work this paper cites.
1d convolutional neural networks and applications: A survey, 2019
Kiranyaz, S., Avci, O., Abdeljaber, O., Ince, T., Gabbouj, M., and Inman, D. J · 1905
Earlier work this paper cites.
Mutual information scaling and expressive power of sequence models, 2019
Shen, H · 1905
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 1905
Earlier work this paper cites.
A multiscale visualization of attention in the transformer model, 2019
Vig, J · 1906
Earlier work this paper cites.
Tsai, Y.-H. H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdinov, R · 1908
Earlier work this paper cites.
Effective construction of linear state-variable models from input/output functions
Ho, B. L. and Kalman, R. E · 1966
Earlier work this paper cites.
Stochastic theory of minimal realization
Akaike, H · 1974
Earlier work this paper cites.
On the stochastic realization problem
Lindquist, A. and Picci, G · 1979
Earlier work this paper cites.
Learning Internal Representations by Error Propagation , pp. 318–362
Rumelhart, D. E. and McClelland, J. L · 1987
Earlier work this paper cites.
Models for Dynamics , pp. 171–269
Willems, J. C · 1989
Earlier work this paper cites.
Model reduction in limited time and frequency intervals
Gawronski, W. and Juang, J.-N · 1990
Earlier work this paper cites.
Approximation of fir by iir digital filters: An algorithm based on balanced model reduction
Beliczynski, B., Kale, I., and Cain, G. D · 1992
Earlier work this paper cites.
Introductory digital signal processing with computer applications (revised ed.)
Lynn, P. A. and Fuerst, W · 1994
Earlier work this paper cites.
Signals & systems (2nd ed.)
Oppenheim, A. V., Willsky, A. S., and Nawab, S. H · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Linear System Theory and Design
Chen, C.-T · 1998
Earlier work this paper cites.
Time-Varying Systems and Computations
DeWilde, P. and van der Veen, A · 1998
Earlier work this paper cites.
System Identification: Theory for the User
Ljung, L · 1999
Earlier work this paper cites.
Low-rank bottleneck in multi-head attention models, 2020
Bhojanapalli, S., Yun, C., Rawat, A. S., Reddi, S. J., and Kumar, S · 2002
Earlier work this paper cites.
Massaroli, S., Poli, M., Park, J., Yamashita, A., and Asama, H · 2002
Earlier work this paper cites.
Glu variants improve transformer, 2020
Shazeer, N · 2002
Earlier work this paper cites.
Quantifying attention flow in transformers, 2020
Abnar, S. and Zuidema, W · 2005
Earlier work this paper cites.
A note on the representation and definition of semiseparable matrices
Vandebril, R., Van Barel, M., and Mastronardi, N · 2005
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention, 2020
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2006
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality
Roy, O. and Vetterli, M · 2007
Cited alongside, same era.
Hopfield networks is all you need, 2021
Ramsauer, H., Schäfl, B., Lehner, J., Seidl, P., Widrich, M., Adler, T., Gruber, L., Holzleitner, M., Pavlović, M., Sandve, G. K., Greiff, V., Kreil, D., Kopp, M., Klambauer, G., Brandstetter, J., and Hochreiter, S · 2008
Cited alongside, same era.
Measuring massive multitask language understanding, 2021
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2009
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Cited alongside, same era.
Fast algorithms for hierarchically semiseparable matrices
Xia, J., Chandrasekaran, S., Gu, M., and Li, X. S · 2010
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding, 2023
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models, 2023
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F · 2023
Later among the works it cites.
Leveraging low-rank and sparse recurrent connectivity for robust closed-loop control, 2023
Tumma, N., Lechner, M., Loo, N., Hasani, R., and Rus, D · 2023
Later among the works it cites.
Attention is all you need, 2023
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2023
Later among the works it cites.
In-context language learning: Architectures and algorithms, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the difficulty of training recurrent neural networks, 2013
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Cited alongside, same era.
Using fast weights to attend to the recent past, 2016
Ba, J., Hinton, G., Mnih, V., Leibo, J. Z., and Ionescu, C · 2016
Cited alongside, same era.
Parallelizing linear recurrent neural nets over sequence length, 2018
Martin, E. and Cundy, C · 2018
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Thread: Circuits
Cammarata, N., Carter, S., Goh, G., Olah, C., Petrov, M., Schubert, L., Voss, C., Egan, B., and Lim, S. K · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
Liquid structural state-space models, 2022
Hasani, R., Lechner, M., Wang, T.-H., Chahine, M., Amini, A., and Rus, D · 2022
Cited alongside, same era.
Akyürek, E., Wang, B., Kim, Y., and Andreas, J · 2024
Later among the works it cites.
The hidden attention of mamba models, 2024
Ali, A., Zimerman, I., and Wolf, L · 2024
Later among the works it cites.
Physics of language models: Part 3.3, knowledge capacity scaling laws, 2024
Allen-Zhu, Z. and Li, Y · 2024
Later among the works it cites.
When benchmarks are targets: Revealing the sensitivity of large language model leaderboards, 2024
Alzahrani, N., Alyahya, H. A., Alnumay, Y., Alrashed, S., Alsubaie, S., Almushaykeh, Y., Mirza, F., Alotaibi, N., Altwairesh, N., Alowisheq, A., Bari, M. S., and Khan, H · 2024
Later among the works it cites.
Simple linear attention language models balance the recall-throughput tradeoff, 2024
Arora, S., Eyuboglu, S., Zhang, M., Timalsina, A., Alberti, S., Zinsley, D., Zou, J., Rudra, A., and Ré, C · 2024
Later among the works it cites.
Transformers to ssms: Distilling quadratic knowledge to subquadratic models, 2024
Bick, A., Li, K. Y., Xing, E. P., Kolter, J. Z., and Gu, A · 2024
Later among the works it cites.
Dao, T. and Gu, A · 2024
Later among the works it cites.
Griffin: Mixing gated linear recurrences with local attention for efficient language models, 2024
De, S., Smith, S. L., Fernando, A., Botev, A., Cristian-Muraru, G., Gu, A., Haroun, R., Berrada, L., Chen, Y., Srinivasan, S., Desjardins, G., Doucet, A., Budden, D., Teh, Y. W., Pascanu, R., Freitas, N. D., and Gulcehre, C · 2024
Later among the works it cites.
Zamba: A compact 7b ssm hybrid model, 2024
Glorioso, P., Anthony, Q., Tokpanov, Y., Whittington, J., Pilault, J., Ibrahim, A., and Millidge, B · 2024
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces, 2024
Gu, A. and Dao, T · 2024
Later among the works it cites.
Jamba: A hybrid transformer-mamba language model, 2024
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., Abend, O., Alon, R., Asida, T., Bergman, A., Glozman, R., Gokhman, M., Manevich, A., Ratner, N., Rozen, N., Shwartz, E., Zusman, M., and Shoham, Y · 2024
Later among the works it cites.
On the efficiency of transformers: The effect of attention rank, 2024
Min, Z. and Li, Z · 2024
Later among the works it cites.
State-free inference of state-space models: The transfer function approach, 2024
Parnichkun, R. N., Massaroli, S., Moro, A., Smith, J. T. H., Hasani, R., Lechner, M., An, Q., Ré, C., Asama, H., Ermon, S., Suzuki, T., Yamashita, A., and Poli, M · 2024
Later among the works it cites.
The fineweb datasets: Decanting the web for the finest text data at scale, 2024
Penedo, G., Kydlíček, H., allal, L. B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L. V., and Wolf, T · 2024
Later among the works it cites.
Mechanistic design and scaling of hybrid architectures, 2024
Poli, M., Thomas, A. W., Nguyen, E., Ponnusamy, P., Deiseroth, B., Kersting, K., Suzuki, T., Hie, B., Ermon, S., Ré, C., Zhang, C., and Massaroli, S · 2024
Later among the works it cites.
Massive activations in large language models, 2024
Sun, M., Chen, X., Kolter, J. Z., and Liu, Z · 2024
Later among the works it cites.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark, 2024
Wang, Y., Ma, X., Zhang, G., Ni, Y., Chandra, A., Guo, S., Ren, W., Arulraj, A., He, X., Jiang, Z., Li, T., Ku, M., Wang, K., Zhuang, A., Fan, R., Yue, X., and Chen, W · 2024
Later among the works it cites.
On the role of attention masks and layernorm in transformers, 2024
Wu, X., Ajorlou, A., Wang, Y., Jegelka, S., and Jadbabaie, A · 2024
Later among the works it cites.
Efficient streaming language models with attention sinks, 2024
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Later among the works it cites.
B’mojo: Hybrid state space realizations of foundation models with eidetic and fading memory, 2024
Zancato, L., Seshadri, A., Dukler, Y., Golatkar, A., Shen, Y., Bowman, B., Trager, M., Achille, A., and Soatto, S · 2024
Later among the works it cites.