Fetching the paper…
Reading the bibliography…
Linear RNN architectures, like Mamba, can be competitive with Transformer models in language modeling while having advantageous deployment characteristics.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 1905
Earlier work this paper cites.
PubMedQA: A Dataset for Biomedical Research Question Answering
Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., and Lu, X. (2019) · 1909
Earlier work this paper cites.
Pre-trained summarization distillation
Shleifer, S. and Rush, A. M. (2020) · 2010
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Kim, Y. and Rush, A. M. (2016) · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R. (2016) · 2016
Earlier work this paper cites.
RACE: Large-scale ReAding Comprehension Dataset From Examinations
Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. (2018) · 2018
Earlier work this paper cites.
PIQA: Reasoning about Physical Commonsense in Natural Language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al. (2020) · 2020
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2020) · 2020
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D. (2020) · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. (2021) · 2021
Earlier work this paper cites.
Efficiently Modeling Long Sequences with Structured State Spaces
Gu, A., Goel, K., and Ré, C. (2021) · 2021
Earlier work this paper cites.
Longt5: Efficient text-to-text transformer for long sequences
Guo, M., Ainslie, J., Uthus, D., Ontanon, S., Ni, J., Sung, Y.-H., and Yang, Y. (2021) · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2021) · 2021
Earlier work this paper cites.
Going beyond linear transformers with recurrent fast weight programmers
Irie, K., Schlag, I., Csordás, R., and Schmidhuber, J. (2021) · 2021
Earlier work this paper cites.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Lin, S., Hilton, J., and Evans, O. (2021) · 2021
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. (2021) · 2021
Earlier work this paper cites.
Linear transformers are secretly fast weight programmers
Schlag, I., Irie, K., and Schmidhuber, J. (2021) · 2021
Earlier work this paper cites.
Hungry hungry hippos: Towards language modeling with state space models
Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A., and Re, C. (2022) · 2022
Earlier work this paper cites.
On the Parameterization and Initialization of Diagonal State Space Models
Gu, A., Goel, K., Gupta, A., and Ré, C. (2022) · 2022
Earlier work this paper cites.
Diagonal State Spaces are as Effective as Structured State Spaces
Gupta, A., Gu, A., and Berant, J. (2022) · 2022
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O. (2022) · 2022
Cited alongside, same era.
Scrolls: Standardized comparison over long language sequences
Shaham, U., Segal, E., Ivgi, M., Efrat, A., Yoran, O., Haviv, A., Gupta, A., Xiong, W., Geva, M., Berant, J., et al. (2022) · 2022
Cited alongside, same era.
Wang, J., Yan, J. N., Gu, A., and Rush, A. M. (2022) · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2022) · 2022
Cited alongside, same era.
Simple linear attention language models balance the recall-throughput tradeoff
Arora, S., Eyuboglu, S., Zhang, M., Timalsina, A., Alberti, S., Zinsley, D., Zou, J., Rudra, A., and Ré, C. (2024) · 2024
Closest in time.
Infinity instruct
BAAI (2024) · 2024
Closest in time.
xlstm: Extended long short-term memory
Beck, M., Pöppel, K., Spanring, M., Auer, A., Prudnikova, O., Kopp, M., Klambauer, G., Brandstetter, J., and Hochreiter, S. (2024) · 2024
Closest in time.
Speculative streaming: Fast llm inference without auxiliary models
Bhendawade, N., Belousova, I., Fu, Q., Mason, H., Rastegari, M., and Najibi, M. (2024) · 2024
Closest in time.
Transformers to ssms: Distilling quadratic knowledge to subquadratic models
Bick, A., Li, K. Y., Xing, E. P., Kolter, J. Z., and Gu, A. (2024) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zoology: Measuring and improving recall in efficient language models
Arora, S., Eyuboglu, S., Timalsina, A., Johnson, I., Poli, M., Zou, J., Rudra, A., and Ré, C. (2023) · 2023
Cited alongside, same era.
Ultrafeedback: Boosting language models with high-quality feedback
Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M. (2023) · 2023
Cited alongside, same era.
Enhancing chat language models by scaling high-quality instructional conversations
Ding, N., Chen, Y., Xu, B., Qin, Y., Hu, S., Liu, Z., Sun, M., and Zhou, B. (2023) · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. (2023) · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T. (2023) · 2023
Cited alongside, same era.
Rest: Retrieval-based speculative decoding
He, Z., Zhong, Z., Cai, T., Lee, J. D., and He, D. (2023) · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. (2023) · 2023
Cited alongside, same era.
Closest in time.
Recurrentgemma: Moving past transformers for efficient open language models
Botev, A., De, S., Smith, S. L., Fernando, A., Muraru, G.-C., Haroun, R., Berrada, L., Pascanu, R., Sessa, P. G., Dadashi, R., et al. (2024) · 2024
Closest in time.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., Lee, J. D., Chen, D., and Dao, T. (2024) · 2024
Closest in time.
Genqa: Generating millions of instructions from a handful of prompts
Chen, J., Qadri, R., Wen, Y., Jain, N., Kirchenbauer, J., Zhou, T., and Goldstein, T. (2024) · 2024
Closest in time.
Dao, T. and Gu, A. (2024) · 2024
Closest in time.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
De, S., Smith, S. L., Fernando, A., Botev, A., Cristian-Muraru, G., Gu, A., Haroun, R., Berrada, L., Chen, Y., Srinivasan, S., et al. (2024) · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. (2024) · 2024
Closest in time.
Cruxeval: A benchmark for code reasoning, understanding and execution
Gu, A., Rozière, B., Leather, H., Solar-Lezama, A., Synnaeve, G., and Wang, S. I. (2024) · 2024
Closest in time.
Jamba: A hybrid transformer-mamba language model
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., et al. (2024) · 2024
Closest in time.
ZeroEval: A Unified Framework for Evaluating Language Models
Lin, B. Y. (2024) · 2024
Closest in time.
Laughing hyena distillery: Extracting compact recurrences from convolutions
Massaroli, S., Poli, M., Fu, D., Kumbong, H., Parnichkun, R., Romero, D., Timalsina, A., McIntyre, Q., Chen, B., Rudra, A., et al. (2024) · 2024
Closest in time.
Simpo: Simple preference optimization with a reference-free reward
Meng, Y., Xia, M., and Chen, D. (2024) · 2024
Closest in time.
Linearizing large language models
Mercat, J., Vasiljevic, I., Keh, S., Arora, K., Dave, A., Gaidon, A., and Kollar, T. (2024) · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C. (2024) · 2024
Closest in time.
Scavenging hyena: Distilling transformers into long convolution models
Ralambomihanta, T. R., Mohammadzadeh, S., Islam, M. S. N., Jabbour, W., and Liang, L. (2024) · 2024
Closest in time.
Samba: Simple hybrid state space models for efficient unlimited context language modeling
Ren, L., Liu, Y., Lu, Y., Shen, Y., Liang, C., and Chen, W. (2024) · 2024
Closest in time.
An empirical study of mamba-based language models
Waleffe, R., Byeon, W., Riach, D., Norick, B., Korthikanti, V., Dao, T., Gu, A., Hatamizadeh, A., Singh, S., Narayanan, D., et al. (2024) · 2024
Closest in time.
Snakes and ladders: Accelerating SSM inference with speculative decoding
Wu, Y., Dukler, Y., Trager, M., Achille, A., Xia, W., and Soatto, S. (2024) · 2024
Closest in time.
The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry
Zhang, M., Bhatia, K., Kumbong, H., and Ré, C. (2024) · 2024
Closest in time.