Fetching the paper…
Reading the bibliography…
Token-free language models learn directly from raw bytes and remove the inductive bias of subword tokenization.
Prefix Sums and Their Applications
Guy E Blelloch · 1990
Earlier work this paper cites.
A Neural Probabilistic Language Model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2000
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
Taku Kudo and John Richardson · 2012
Earlier work this paper cites.
Japanese and Korean Voice Search
Mike Schuster and Kaisuke Nakajima · 2012
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Earlier work this paper cites.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
A Simple Method for Commonsense Reasoning
Trieu H. Trinh and Quoc V. Le · 2018
Earlier work this paper cites.
Character-Level Language Modeling with Deeper Self-Attention
Rami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, and Llion Jones · 2019
Earlier work this paper cites.
Bridging the gap for tokenizer-free language models
Dokook Choe, Rami Al-Rfou, Mandy Guo, Heeyoung Lee, and Noah Constant · 2019
Earlier work this paper cites.
Spell once, summon anywhere: A two-level open-vocabulary language model
Sebastian J Mielke and Jason Eisner · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing
Zihang Dai, Guokun Lai, Yiming Yang, and Quoc Le · 2020
Earlier work this paper cites.
The Curious Case of Neural Text Degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Earlier work this paper cites.
Charbert: character-aware pre-trained language model
Wentao Ma, Yiming Cui, Chenglei Si, Ting Liu, Shijin Wang, and Guoping Hu · 2020
Earlier work this paper cites.
Compressive Transformers for Long-Range Sequence Modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap · 2020
Earlier work this paper cites.
Neural Machine Translation with Byte-Level Subwords
Changhan Wang, Kyunghyun Cho, and Jiatao Gu · 2020
Cited alongside, same era.
Efficiently Modeling Long Sequences with Structured State Spaces
Albert Gu, Karan Goel, and Christopher Ré · 2021
Cited alongside, same era.
Efficient Content-Based Sparse Attention with Routing Transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Cited alongside, same era.
Bygpt5: End-to-end style-conditioned poetry generation with token-free language models
Jonas Belouadi and Steffen Eger · 2022
Cited alongside, same era.
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Albert Gu and Tri Dao · 2023
Later among the works it cites.
How to Train your HiPPO: State Space Models with Generalized Orthogonal Basis Projections
Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher Re · 2023
Later among the works it cites.
Rest: Retrieval-based speculative decoding
Zhenyu He, Zexuan Zhong, Tianle Cai, Jason D Lee, and Di He · 2023
Later among the works it cites.
Fast Inference from Transformers via Speculative Decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Later among the works it cites.
Xiaoxuan Liu, Lanxiang Hu, Peter Bailis, Ion Stoica, Zhijie Deng, Alvin Cheung, and Hao Zhang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Canine: Pre-training an Efficient Tokenization-Free Encoder for Language Representation
Jonathan H Clark, Dan Garrette, Iulia Turc, and John Wieting · 2022
Cited alongside, same era.
Hungry hungry hippos: Towards language modeling with state space models
Daniel Y Fu, Tri Dao, Khaled Kamal Saab, Armin W Thomas, Atri Rudra, and Christopher Re · 2022
Cited alongside, same era.
It’s raw! audio generation with state-space models
Karan Goel, Albert Gu, Chris Donahue, and Christopher Ré · 2022
Cited alongside, same era.
On the Parameterization and Initialization of Diagonal State Space Models
Albert Gu, Karan Goel, Ankit Gupta, and Christopher Ré · 2022
Cited alongside, same era.
Diagonal State Spaces are as Effective as Structured State Spaces
Ankit Gupta, Albert Gu, and Jonathan Berant · 2022
Cited alongside, same era.
General-purpose, long-context autoregressive modeling with Perceiver AR
Curtis Hawthorne, Andrew Jaegle, Cătălina Cangea, Sebastian Borgeaud, Charlie Nash, Mateusz Malinowski, Sander Dieleman, Oriol Vinyals, Matthew Botvinick, Ian Simon, Hannah Sheahan, Neil Zeghidour, Jean-Baptiste Alayrac, Joao Carreira, and Jesse Engel · 2022
Cited alongside, same era.
Block-Recurrent Transformers
DeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer, and Behnam Neyshabur · 2022
Cited alongside, same era.
Long Range Language Modeling via Gated State Spaces
Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur · 2023
Later among the works it cites.
Resurrecting Recurrent Neural Networks for Long Sequences
Antonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De · 2023
Later among the works it cites.
Simplified State Space Layers for Sequence Modeling
Jimmy T.H. Smith, Andrew Warrington, and Scott Linderman · 2023
Later among the works it cites.
Accelerating llm inference with staged speculative decoding
Benjamin Spector and Chris Re · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation
Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui · 2023
Later among the works it cites.
MegaByte: Predicting Million-byte Sequences with Multiscale Transformers
Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis · 2023
Later among the works it cites.
Simple linear attention language models balance the recall-throughput tradeoff
Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, and Christopher Ré · 2024
Closest in time.
Speculative streaming: Fast llm inference without auxiliary models
Nikhil Bhendawade, Irina Belousova, Qichen Fu, Henry Mason, Mohammad Rastegari, and Mahyar Najibi · 2024
Closest in time.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
Soham De, Samuel L Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, et al · 2024
Closest in time.
Caduceus: Bi-directional equivariant long-range dna sequence modeling
Yair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao, Albert Gu, and Volodymyr Kuleshov · 2024
Closest in time.
Focused transformer: Contrastive training for context scaling
Szymon Tworkowski, Konrad Staniszewski, Mikołaj Pacek, Yuhuai Wu, Henryk Michalewski, and Piotr Miłoś · 2024
Closest in time.
Diffusion models without attention
Jing Nathan Yan, Jiatao Gu, and Alexander M Rush · 2024
Closest in time.