Fetching the paper…
Reading the bibliography…
Bytes form the basis of the digital world and thus are a promising building block for multimodal foundation models.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 1911
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2012
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C. L., and Girshick, R. B · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Yi, K., Wu, J., Gan, C., Torralba, A., Kohli, P., and Tenenbaum, J. B · 2018
Earlier work this paper cites.
Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes
Li, B., Zhang, Y., Sainath, T., Wu, Y., and Chan, W · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
Neural machine translation with byte-level subwords
Wang, C., Cho, K., and Gu, J · 2020
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P.-S., et al · 2021
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
Gu, A., Goel, K., and Re, C · 2022
Earlier work this paper cites.
Perceiver IO: A general architecture for structured inputs & outputs
Jaegle, A., Borgeaud, S., Alayrac, J.-B., Doersch, C., Ionescu, C., Ding, D., Koppula, S., Zoran, D., Brock, A., Shelhamer, E., Henaff, O. J., Botvinick, M., Zisserman, A., Vinyals, O., and Carreira, J · 2022
Earlier work this paper cites.
ByT5: Towards a token-free future with pre-trained byte-to-byte models
Xue, L., Barua, A., Constant, N., Al-Rfou, R., Narang, S., Kale, M., Roberts, A., and Raffel, C · 2022
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2023
Cited alongside, same era.
On the computational complexity of self-attention
Keles, F., Wijewardena, P., and Hegde, C · 2023
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., et al · 2023
Cited alongside, same era.
The llama 3 herd of models, 2024
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., Rao, A., Zhang, A., Rodriguez, A., et al · 2024
Later among the works it cites.
On the parameterization and initialization of diagonal state space models
Gu, A., Gupta, A., Goel, K., and Ré, C · 2024
Later among the works it cites.
Parallel prefix sum (scan) with cuda, 2024
Harris, M., Sengupta, S., and Owens, J. D · 2024
Later among the works it cites.
Bytes are all you need: Transformers operating directly on file bytes
Horton, M., Mehta, S., Farhadi, A., and Rastegari, M · 2024
Later among the works it cites.
Jamba: A hybrid transformer-mamba language model, 2024
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., Abend, O., Alon, R., Asida, T., Bergman, A., Glozman, R., Gokhman, M., Manevich, A., Ratner, N., Rozen, N., Shwartz, E., Zusman, M., and Shoham, Y · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Megabyte: Predicting million-byte sequences with multiscale transformers
Yu, L., Simig, D., Flaherty, C., Aghajanyan, A., Zettlemoyer, L., and Lewis, M · 2023
Cited alongside, same era.
Blackmamba: Mixture of experts for state-space models
Anthony, Q. G., Tokpanov, Y., Glorioso, P., and Millidge, B · 2024
Cited alongside, same era.
Decimamba: Exploring the length extrapolation potential of mamba, 2024
Ben-Kish, A., Zimerman, I., Abu-Hussein, S., Cohen, N., Globerson, A., Wolf, L., and Giryes, R · 2024
Cited alongside, same era.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2024
Cited alongside, same era.
Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality
Dao, T. and Gu, A · 2024
Cited alongside, same era.
Zamba: A compact 7b ssm hybrid model, 2024
Glorioso, P., Anthony, Q., Tokpanov, Y., Whittington, J., Pilault, J., Ibrahim, A., and Millidge, B · 2024
Cited alongside, same era.
Later among the works it cites.
Learn to explain: multimodal reasoning via thought chains for science question answering
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A · 2024
Later among the works it cites.
Byte latent transformer: Patches scale better than tokens, 2024
Pagnoni, A., Pasunuru, R., Rodriguez, P., Nguyen, J., Muller, B., Li, M., Zhou, C., Yu, L., Weston, J., Zettlemoyer, L., Ghosh, G., Lewis, M., Holtzman, A., and Iyer, S · 2024
Later among the works it cites.
An empirical study of mamba-based language models
Waleffe, R., Byeon, W., Riach, D., Norick, B., Korthikanti, V., Dao, T., Gu, A., Hatamizadeh, A., Singh, S., Narayanan, D., Kulshreshtha, G., Singh, V., Casper, J., Kautz, J., Shoeybi, M., and Catanzaro, B · 2024
Later among the works it cites.
Mambabyte: Token-free selective state space model
Wang, J., Gangavarapu, T., Yan, J. N., and Rush, A. M · 2024
Later among the works it cites.
Beyond language models: Byte models are digital world simulators, 2024
Wu, S., Tan, X., Wang, Z., Wang, R., Li, X., and Sun, M · 2024
Later among the works it cites.
Length extrapolation of transformers: A survey from the perspective of positional encoding
Zhao, L., Feng, X., Feng, X., Zhong, W., Xu, D., Yang, Q., Liu, H., Qin, B., and Liu, T · 2024
Later among the works it cites.