Fetching the paper…
Reading the bibliography…
Natural language is composed of words, but modern large language models (LLMs) process sub-words as input.
A generalized solution of the orthogonal procrustes problem
Peter H Schönemann · 1966
Earlier work this paper cites.
Morphology and meaning in the english mental lexicon
William Marslen-Wilson, Lorraine K Tyler, Rachelle Waksler, and Lianne Older · 1994
Earlier work this paper cites.
Perception of wordlikeness: Effects of segment probability and length on the processing of nonwords
Stefan A. Frisch, Nathan R. Large, and David B. Pisoni · 2000
Earlier work this paper cites.
358,534 nonwords: the arc nonword database
Kathleen Rastle, Jonathan Harrington, and Max Coltheart · 2002
Earlier work this paper cites.
Words in the mind: An introduction to the mental lexicon
Jean Aitchison · 2012
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni · 2018
Earlier work this paper cites.
Martin Gerlach and Francesc Font-Clos · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo · 2018
Earlier work this paper cites.
Adaptively sparse transformers
Gonçalo M. Correia, Vlad Niculae, and André F. T. Martins · 2019
Earlier work this paper cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
Root mean square layer normalization
Biao Zhang and Rico Sennrich · 2019
Earlier work this paper cites.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett · 2020
Earlier work this paper cites.
Emerging trends: Subwords, seriously?
Kenneth Ward Church · 2020
Earlier work this paper cites.
Analyzing individual neurons in pre-trained language models
Nadir Durrani, Hassan Sajjad, Fahim Dalvi, and Yonatan Belinkov · 2020
Earlier work this paper cites.
Wiki-40B: Multilingual language model dataset
Mandy Guo, Zihang Dai, Denny Vrandečić, and Rami Al-Rfou · 2020
Earlier work this paper cites.
Don‘t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith · 2020
Earlier work this paper cites.
Getting the## life out of living: How adequate are word-pieces for modelling complex morphology?
Stav Klein and Reut Tsarfaty · 2020
Earlier work this paper cites.
interpreting GPT: the logit lens
nostalgebraist · 2020
Earlier work this paper cites.
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita · 2020
Earlier work this paper cites.
Probing pretrained language models for lexical semantics
Ivan Vulić, Edoardo Maria Ponti, Robert Litschko, Goran Glavaš, and Anna Korhonen · 2020
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Cited alongside, same era.
Adapting pre-trained language models to African languages via multilingual adaptive fine-tuning
Jesujoba O. Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow · 2022
Cited alongside, same era.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei · 2022
Cited alongside, same era.
Softmax linear units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah · 2022
Cited alongside, same era.
Incorporating context into subword vocabularies
Shaked Yehezkel and Yuval Pinter · 2023
Later among the works it cites.
Tokenization and the noiseless channel
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell · 2023
Later among the works it cites.
Yi: Open foundation models by 01.ai, 2024
01. AI, :, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and Zonghong Dai · 2024
Closest in time.
Evaluating subword tokenization: Alien subword composition and oov generalization challenge
Khuyagbaatar Batsuren, Ekaterina Vylomova, Verna Dankers, Tsetsuukhei Delgerbaatar, Omri Uzan, Yuval Pinter, and Gábor Bella · 2024
Closest in time.
Bpe-knockout: Pruning pre-existing bpe tokenisers with backwards-compatible morphological semi-supervision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fast vocabulary transfer for language model compression
Leonidas Gee, Andrea Zugarini, Leonardo Rigutini, and Paolo Torroni · 2022
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg · 2022
Cited alongside, same era.
An embarrassingly simple method to mitigate undesirable properties of pretrained language model tokenizers
Valentin Hofmann, Hinrich Schuetze, and Janet Pierrehumbert · 2022
Cited alongside, same era.
What do tokens know about their characters and how do they know it?
Ayush Kaushal and Kyle Mahowald · 2022
Cited alongside, same era.
Analyzing encoded concepts in transformer language models
Hassan Sajjad, Nadir Durrani, Fahim Dalvi, Firoj Alam, Abdul Khan, and Jia Xu · 2022
Cited alongside, same era.
Do all languages cost the same? tokenization in the era of commercial language models
Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David Mortensen, Noah Smith, and Yulia Tsvetkov · 2023
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens, 2023
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Cited alongside, same era.
Thomas Bauwens and Pieter Delobelle · 2024
Closest in time.
Attend first, consolidate later: On the importance of attention in different LLM layers
Amit Ben Artzy and Roy Schwartz · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Information flow routes: Automatically interpreting language models at scale
Javier Ferrando and Elena Voita · 2024
Closest in time.
Token erasure as a footprint of implicit vocabulary items in LLMs
Sheridan Feucht, David Atkinson, Byron C Wallace, and David Bau · 2024
Closest in time.
Patchscopes: A unifying framework for inspecting hidden representations of language models
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva · 2024
Closest in time.
Efficient and effective vocabulary expansion towards multilingual large language models
Seungduk Kim, Seungtaek Choi, and Myeongho Jeong · 2024
Closest in time.
The remarkable robustness of llms: Stages of inference?, 2024
Vedang Lad, Wes Gurnee, and Max Tegmark · 2024
Closest in time.
Language models implement simple Word2Vec-style vector arithmetic
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick · 2024
Closest in time.
The geometry of categorical and hierarchical concepts in large language models
Kiho Park, Yo Joong Choe, Yibo Jiang, and Victor Veitch · 2024
Closest in time.
Tokenization is more than compression
Craig W Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris Tanner · 2024
Closest in time.
Function vectors in large language models
Eric Todd, Millicent Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau · 2024
Closest in time.
Benchmarking retrieval-augmented generation for medicine
Guangzhi Xiong, Qiao Jin, Zhiyong Lu, and Aidong Zhang · 2024
Closest in time.
Investigating continual pretraining in large language models: Insights and implications
Çağatay Yıldız, Nishaanth Kanna Ravichandran, Prishruit Punia, Matthias Bethge, and Beyza Ermis · 2024
Closest in time.
Jump to conclusions: Short-cutting transformers with linear transformations
Alexander Yom Din, Taelin Karidi, Leshem Choshen, and Mor Geva · 2024
Closest in time.
Towards best practices of activation patching in language models: Metrics and methods
Fred Zhang and Neel Nanda · 2024
Closest in time.
Llama beyond english: An empirical study on language capability transfer
Jun Zhao, Zhihao Zhang, Qi Zhang, Tao Gui, and Xuanjing Huang · 2024
Closest in time.
The representation geometry of features and hierarchy in large language models
Kiho Park, Yo Joong Choe, Yibo Jiang, and Victor Veitch · 2025
Closest in time.