Fetching the paper…
Reading the bibliography…
Modern recurrent layers are emerging as a promising path toward edge deployment of foundation models, especially in the context of large language models (LLMs).
HellaSwag: Can a Machine Really Finish Your Sentence?, May 2019
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 1905
Earlier work this paper cites.
WinoGrande: An Adversarial Winograd Schema Challenge at Scale, November 2019
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 1907
Earlier work this paper cites.
PIQA: Reasoning about Physical Commonsense in Natural Language, November 2019
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 1911
Earlier work this paper cites.
Language Models are Few-Shot Learners, July 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
HiPPO: Recurrent Memory with Optimal Polynomial Projections, October 2020
Gu, A., Dao, T., Ermon, S., Rudra, A., and Re, C · 2008
Earlier work this paper cites.
Pointer Sentinel Mixture Models, September 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context, June 2016
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D · 2017
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding, February 2019
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Earlier work this paper cites.
Hardware aware training for efficient keyword spotting on general purpose and specialized hardware
Blouw, P., Malik, G., Morcos, B., Voelker, A., and Eliasmith, C · 2021
Cited alongside, same era.
Understanding and Overcoming the Challenges of Efficient Transformer Quantization, September 2021
Bondarenko, Y., Nagel, M., and Blankevoort, T · 2021
Cited alongside, same era.
Do long-range language models actually use long-range context?, 2021
Sun, S., Krishna, K., Mattarella-Micke, A., and Iyyer, M · 2021
Cited alongside, same era.
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, November 2022
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Cited alongside, same era.
Efficiently Modeling Long Sequences with Structured State Spaces, August 2022
Gu, A., Goel, K., and Ré, C · 2022
Attention Is All You Need, August 2023
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2023
Later among the works it cites.
RecurrentGemma: Moving Past Transformers for Efficient Open Language Models, April 2024
Botev, A., De, S., Smith, S. L., Fernando, A., Muraru, G.-C., Haroun, R., Berrada, L., Pascanu, R., Sessa, P. G., Dadashi, R., Hussenot, L., Ferret, J., Girgin, S., Bachem, O., Andreev, A., Kenealy, K., Mesnard, T., Hardin, C., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., Tafti, P., Joulin, A., Fiedel, N., Senter, E., Chen, Y., Srinivasan, S., Desjardins, G., Budden, D., Doucet, A., Vikram, S., Paszke, A., Gale, T., Borgeaud, S., Chen, C., Brock, A., Paterson, A., Brennan, J., Risdal, M., Gundluru, R., Devanathan, N., Mooney, P., Chauhan, N., Culliton, P., Martins, L. G., Bandy, E., Huntsperger, D., Cameron, G., Zucker, A., Warkentin, T., Peran, L., Giang, M., Ghahramani, Z., Farabet, C., Kavukcuoglu, K., Hassabis, D., Hadsell, R., Teh, Y. W., and de Frietas, N · 2024
Closest in time.
Transformers are ssms: Generalized models and efficient algorithms through structured state space duality, 2024
Dao, T. and Gu, A · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing, November 2023
Bondarenko, Y., Nagel, M., and Blankevoort, T · 2023
Cited alongside, same era.
Mamba: Linear-Time Sequence Modeling with Selective State Spaces, December 2023
Gu, A. and Dao, T · 2023
Cited alongside, same era.
RWKV: Reinventing RNNs for the Transformer Era, December 2023
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., He, X., Hou, H., Lin, J., Kazienko, P., Kocon, J., Kong, J., Koptyra, B., Lau, H., Mantri, K. S. I., Mom, F., Saito, A., Song, G., Tang, X., Wang, B., Wind, J. S., Wozniak, S., Zhang, R., Zhang, Z., Zhao, Q., Zhou, P., Zhou, Q., Zhu, J., and Zhu, R.-J · 2023
Cited alongside, same era.
Hyena Hierarchy: Towards Larger Convolutional Language Models, April 2023
Poli, M., Massaroli, S., Nguyen, E., Fu, D. Y., Dao, T., Baccus, S., Bengio, Y., Ermon, S., and Ré, C · 2023
Cited alongside, same era.
Choice of Plausible Alternatives: An Evaluation of Commonsense Causal Reasoning
Roemmele, M., Bejan, C. A., and Gordon, A. S
Cited in the paper.
De, S., Smith, S. L., Fernando, A., Botev, A., Cristian-Muraru, G., Gu, A., Haroun, R., Berrada, L., Chen, Y., Srinivasan, S., Desjardins, G., Doucet, A., Budden, D., Teh, Y. W., Pascanu, R., De Freitas, N., and Gulcehre, C · 2024
Closest in time.
Jamba: A Hybrid Transformer-Mamba Language Model, March 2024
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., Abend, O., Alon, R., Asida, T., Bergman, A., Glozman, R., Gokhman, M., Manevich, A., Ratner, N., Rozen, N., Shwartz, E., Zusman, M., and Shoham, Y · 2024
Closest in time.
Efficient Video and Audio Processing with Loihi 2
Shrestha, S. B., Timcheck, J., Frady, P., Campos-Macias, L., and Davies, M · 2024
Closest in time.
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, March 2024
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2024
Closest in time.
Zhang, Y., Yang, F., Peng, S., Wang, F., and Pan, A · 2024
Closest in time.