Fetching the paper…
Reading the bibliography…
Tokenizer transplant in cross-vocabulary model composition reconstructs donor-only embedding rows as weighted combinations over shared lexical anchors and reuses those coefficients on the base.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul F. Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
Remarques sur un résultat non publié de B. Maurey
Gilles Pisier · 1981
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R. Barron · 1993
Earlier work this paper cites.
Orthogonal matching pursuit: Recursive function approximation with applications to wavelet decomposition
Yagyensh Chandra Pati, Ramin Rezaiifar, and Perinkulam Sambamurthy Krishnaprasad · 1993
Earlier work this paper cites.
Greed is good: Algorithmic results for sparse approximation
Joel A. Tropp · 2004
Earlier work this paper cites.
Stable recovery of sparse overcomplete representations in the presence of noise
David L. Donoho, Michael Elad, and Vladimir N. Temlyakov · 2006
Earlier work this paper cites.
Signal recovery from random measurements via orthogonal matching pursuit
Joel A. Tropp and Anna C. Gilbert · 2007
Earlier work this paper cites.
Measuring massive multitask language understanding, 2020
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2009
Earlier work this paper cites.
Smallest singular value of a random rectangular matrix
Mark Rudelson and Roman Vershynin · 2009
Earlier work this paper cites.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano · 2009
Earlier work this paper cites.
Orthogonal matching pursuit for sparse signal recovery with noise
T. Tony Cai and Lie Wang · 2011
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernandez · 2016
Earlier work this paper cites.
BadNets: Identifying vulnerabilities in the machine learning model supply chain
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Trojaning attack on neural networks
Yingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee, Juan Zhai, Weihang Wang, and Xiangyu Zhang · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Privacy risks of general-purpose language models
X. Pan, M. Zhang, S. Ji, and M. Yang · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Probing toxic content in large pre-trained language models
Nedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song, and Dit-Yan Yeung · 2021
Earlier work this paper cites.
UNKs everywhere: Adapting multilingual language models to new scripts
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder · 2021
Earlier work this paper cites.
How good is your tokenizer? on the monolingual performance of multilingual language models
P. Rust, J. Pfeiffer, I. Vulić, S. Ruder, and I. Gurevych · 2021
Earlier work this paper cites.
Subword mapping and anchoring across languages
Giorgos Vernikos and Andrei Popescu-Belis · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2021
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, et al · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
Merging models with fisher-weighted averaging
Michael Matena and Colin Raffel · 2022
Earlier work this paper cites.
Wechsel: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models
Benjamin Minixhofer, Fabian Paischer, and Navid Rekabsaz · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe · 2022
Cited alongside, same era.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Cited alongside, same era.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt · 2022
Cited alongside, same era.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han · 2022
Cited alongside, same era.
Fishing for magikarp: Automatically detecting under-trained tokens in large language models, 2024
Sander Land and Max Bartolo · 2024
Later among the works it cites.
Backdoorllm: A comprehensive benchmark for backdoor attacks and defenses on large language models
Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, and Jun Sun · 2024
Later among the works it cites.
CROW: Eliminating backdoors from large language models via internal consistency regularization
N. M. Min, L. H. Pham, Y. Li, and J. Sun · 2024
Later among the works it cites.
Benjamin Minixhofer, Edoardo Maria Ponti, and Ivan Vulić · 2024
Later among the works it cites.
Un ministral, des ministraux
Mistral AI team · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Code alpaca: An instruction-following llama model for code generation
Sahil Chaudhary · 2023
Cited alongside, same era.
Focus: Effective embedding initialization for monolingual specialization of multilingual models
Konstantin Dobler and Gerard de Melo · 2023
Cited alongside, same era.
Embedding structure matters: Comparing methods to adapt multilingual vocabularies to new languages
C. M. Downey, Terra Blevins, Nora Goldfine, and Shane Steinert-Threlkeld · 2023
Cited alongside, same era.
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Cited alongside, same era.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2023
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Cited alongside, same era.
Privacy in large language models: Attacks, defenses and future directions
H. Li, Y. Chen, J. Luo, J. Wang, H. Peng, Y. Kang, et al · 2023
Cited alongside, same era.
Efficient language model training through cross-lingual and progressive transfer learning
Malte Ostendorff and Georg Rehm · 2023
Cited alongside, same era.
Later among the works it cites.
An empirical comparison of vocabulary expansion and initialization approaches for language models
Nandini Mundra, Aditya Nanda Kishore Khandavally, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, and Mitesh M. Khapra · 2024
Later among the works it cites.
Jailbreaking and mitigation of vulnerabilities in large language models
Benji Peng, Keyu Chen, Qian Niu, Ziqian Bi, Ming Liu, Pohsun Feng, Tianyang Wang, Lawrence K. Q. Yan, Yizhu Wen, Yichao Zhang, Caitlyn Heqi Yin, and Xinyuan Song · 2024
Later among the works it cites.
François Remy, Pieter Delobelle, Hayastan Avetisyan, Alfiya Khabibullina, Miryam de Lhoneux, and Thomas Demeester · 2024
Later among the works it cites.
Large language model safety: A holistic survey, Dec 2024
Dan Shi, Tianhao Shen, Yufei Huang, Zhigen Li, Yongqi Leng, Renren Jin, Chuang Liu, Xinwei Wu, Zishan Guo, Linhao Yu, Ling Shi, Bojian Jiang, and Deyi Xiong · 2024
Later among the works it cites.
A comprehensive study of jailbreak attack versus defense for large language models
Z. Xu, Y. Liu, G. Deng, Y. Li, and S. Picek · 2024
Later among the works it cites.
An empirical study on cross-lingual vocabulary adaptation for efficient language model inference
Atsuki Yamaguchi, Aline Villavicencio, and Nikolaos Aletras · 2024
Later among the works it cites.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang · 2024
Later among the works it cites.
Jailbreak attacks and defenses against large language models: A survey
Sibo Yi, Yule Liu, Zhen Sun, Tianshuo Cong, Xinlei He, Jiaxing Song, Ke Xu, and Qi Li · 2024
Later among the works it cites.
Smollm2: When smol goes big — data-centric training of a fully open small language model
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíček, Agustín Piqueres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo Larcher, Haojun Zhao, Cyril Zakka, Mathieu Morlon, Colin Raffel, Leandro von Werra, and Thomas Wolf · 2025
Closest in time.
Chen Chen, Xueluan Gong, Ziyao Liu, Weifeng Jiang, Si Qi Goh, and Kwok-Yan Lam · 2025
Closest in time.
Backdoor attacks and countermeasures in natural language processing models: A comprehensive security review
Pengzhou Cheng, Zongru Wu, Wei Du, Haodong Zhao, Wei Lu, and Gongshen Liu · 2025
Closest in time.
Security and privacy challenges of large language models: A survey
B. C. Das, M. H. Amini, and Y. Wu · 2025
Closest in time.
Retrofitting large language models with dynamic tokenization
D. Feher, I. Vulić, and B. Minixhofer · 2025
Closest in time.
Gemma Team · 2025
Closest in time.
Training-free tokenizer transplantation via orthogonal matching pursuit
Charles Goddard and Fernando Fernandes Neto · 2025
Closest in time.
Franken-adapter: Cross-lingual adaptation of LLMs by embedding surgery
F. Jiang, H. Yu, G. Chung, and T. Cohn · 2025
Closest in time.
Semantic aware linear transfer by recycling pre-trained language models for cross-lingual transfer
Seungyoon Lee, Seongtae Hong, Hyeonseok Moon, and Heuiseok Lim · 2025
Closest in time.
Attack and defense techniques in large language models: A survey and new perspectives
Zhiyu Liao, Kang Chen, Yuanguo Lin, Kangkang Li, Yunxuan Liu, Hefeng Chen, Xingwang Huang, and Yuanhui Yu · 2025
Closest in time.
Elba-bench: An efficient learning backdoor attacks benchmark for large language models
Xuxu Liu, Siyuan Liang, Mengya Han, Yong Luo, Aishan Liu, Xiantao Cai, Zheng He, and Dacheng Tao · 2025
Closest in time.
Universal cross-tokenizer distillation via approximate likelihood matching
B. Minixhofer, I. Vulić, and E. M. Ponti · 2025
Closest in time.
Optimizing llms for italian: Reducing token fertility and enhancing efficiency through vocabulary adaptation
Luca Moroni, Giovanni Puccetti, Pere-Lluís Huguet Cabot, Andrei Stefan Bejgu, Alessio Miaschi, Edoardo Barba, Felice Dell’Orletta, Andrea Esuli, and Roberto Navigli · 2025
Closest in time.
Safetyprompts: A systematic review of open datasets for evaluating and improving large language model safety
Paul R"ottger, Fabio Pernisi, Bertie Vidgen, and Dirk Hovy · 2025
Closest in time.
Achieving tokenizer flexibility in language models through heuristic adaptation and supertoken learning, 2025
Shaurya Sharthak, Vinayak Pahalwan, Adithya Kamath, and Adarsh Shirawalmath · 2025
Closest in time.
How can we effectively expand the vocabulary of LLMs with 0.01 GB of target language text?
A. Yamaguchi, A. Villavicencio, and N. Aletras · 2025
Closest in time.
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu · 2025
Closest in time.
Merge hijacking: Backdoor attacks to model merging of large language models
Zenghui Yuan, Yangming Xu, Jiawen Shi, Pan Zhou, and Lichao Sun · 2025
Closest in time.
On large language models safety, security, and privacy: A survey
Ran Zhang, Hong-Wei Li, Xin-Yuan Qian, Wen-Bo Jiang, and Han-Xiao Chen · 2025
Closest in time.
A survey on backdoor threats in large language models (llms): Attacks, defenses, and evaluations
Y. Zhou, T. Ni, W. B. Lee, and Q. Zhao · 2025
Closest in time.