Fetching the paper…
Reading the bibliography…
This paper introduces Uniform Orthogonal Reinitialization Adaptation (UORA), a novel parameter-efficient fine-tuning (PEFT) approach for Large Language Models (LLMs).
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu. 2019 · 1907
Earlier work this paper cites.
LoRAMoE: Alleviating world knowledge forgetting in large language models via MoE-style plugin
Shihan Dou, Enyu Zhou, Yan Liu, Songyang Gao, Wei Shen, Limao Xiong, Yuhao Zhou, Xiao Wang, Zhiheng Xi, Xiaoran Fan, Shiliang Pu, Jiang Zhu, Rui Zheng, Tao Gui, Qi Zhang, and Xuanjing Huang. 2024 · 1945
Earlier work this paper cites.
Exploring versatile generative language model via parameter-efficient transfer learning
Zhaojiang Lin, Andrea Madotto, and Pascale Fung. 2020 · 2004
Earlier work this paper cites.
Adapterfusion: Non-destructive task composition for transfer learning
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2020 · 2005
Earlier work this paper cites.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman. 2008 · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al. 2009 · 2009
Earlier work this paper cites.
Adapterdrop: On the efficiency of adapters in transformers
Andreas Rücklé, Gregor Geigle, Max Glockner, Tilman Beck, Jonas Pfeiffer, Nils Reimers, and Iryna Gurevych. 2020 · 2010
Earlier work this paper cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta. 2020 · 2012
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. 2014 · 2014
Earlier work this paper cites.
Learning to solve arithmetic word problems with verb categorization
Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. 2014 · 2014
Earlier work this paper cites.
Parsing algebraic word problems into equations
Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, and Siena Dumas Ang. 2015 · 2015
Earlier work this paper cites.
Solving general arithmetic word problems
Subhro Roy and Dan Roth. 2015 · 2015
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
Gong Cheng, Junwei Han, and Xiaoqiang Lu. 2017 · 2017
Earlier work this paper cites.
The e2e dataset: New challenges for end-to-end generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2017 · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017 · 2017
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2018 · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang. 2018 · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Training neural networks for and by interpolation
Leonard Berrada, Andrew Zisserman, and M Pawan Kumar. 2020 · 2020
Earlier work this paper cites.
Provable benefit of orthogonal initialization in optimizing deep linear networks
Wei Hu, Lechao Xiao, and Jeffrey Pennington. 2020 · 2020
Earlier work this paper cites.
What’s hidden in a randomly weighted neural network?
Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi, Ali Farhadi, and Mohammad Rastegari. 2020 · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Earlier work this paper cites.
Warp: Word-level adversarial reprogramming
Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021 · 2021
Earlier work this paper cites.
Towards a unified view of parameter-efficient transfer learning
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
On the neural tangent kernel of deep networks with orthogonal initialization
Wei Huang, Weitao Du, and Richard Yi Da Xu. 2021 · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Cited alongside, same era.
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2021 · 2021
Cited alongside, same era.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Elad Ben Zaken, Shauli Ravfogel, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
Revisiting parameter-efficient tuning: Are we really there yet?
Disentangling memory and reasoning ability in large language models
Mingyu Jin, Weidi Luo, Sitao Cheng, Xinyi Wang, Wenyue Hua, Ruixiang Tang, William Yang Wang, and Yongfeng Zhang. 2024 · 2024
Later among the works it cites.
RA-LoRA: Rank-adaptive parameter-efficient fine-tuning for accurate 2-bit quantized large language models
Minsoo Kim, Sihwa Lee, Wonyong Sung, and Jungwook Choi. 2024 · 2024
Later among the works it cites.
Cpseg: Finer-grained image semantic segmentation via chain-of-thought language prompting
Lei Li. 2024 · 2024
Later among the works it cites.
What does the knowledge neuron thesis have to do with knowledge?
Jingcheng Niu, Andrew Liu, Zining Zhu, and Gerald Penn. 2024 · 2024
Later among the works it cites.
Subgraph retrieval enhanced by graph-text alignment for commonsense question answering
Boci Peng, Yongchao Liu, Xiaohe Bo, Sheng Tian, Baokun Wang, Chuntao Hong, and Yan Zhang. 2024a · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guanzheng Chen, Fangyu Liu, Zaiqiao Meng, and Shangsong Liang. 2022 · 2022
Cited alongside, same era.
Does BERT rediscover a classical NLP pipeline?
Jingcheng Niu, Wenjie Lu, and Gerald Penn. 2022 · 2022
Cited alongside, same era.
Mojtaba Valipour, Mehdi Rezagholizadeh, Ivan Kobyzev, and Ali Ghodsi. 2022 · 2022
Cited alongside, same era.
Adamix: Mixture-of-adapter for parameter-efficient tuning of large language models
Yaqing Wang, Subhabrata Mukherjee, Xiaodong Liu, Jing Gao, Ahmed Hassan Awadallah, and Jianfeng Gao. 2022 · 2022
Cited alongside, same era.
When does re-initialization work?
Sheheryar Zaidi, Tudor Berariu, Hyunjik Kim, Jorg Bornschein, Claudia Clopath, Yee Whye Teh, and Razvan Pascanu. 2023 · 2022
Cited alongside, same era.
QLoRA: Efficient finetuning of quantized LLMs
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023 · 2023
Cited alongside, same era.
Sparse low-rank adaptation of pre-trained language models
Ning Ding, Xingtai Lv, Qiaosen Wang, Yulin Chen, Bowen Zhou, Zhiyuan Liu, and Maosong Sun. 2023 · 2023
Cited alongside, same era.
Alignment and generation adapter for efficient video-text understanding
Han Fang, Zhifei Yang, Yuhan Wei, Xianghao Zang, Chao Ban, Zerun Feng, Zhongjiang He, Yongxiang Li, and Hao Sun. 2023 · 2023
Cited alongside, same era.
MELoRA: Mini-ensemble low-rank adapters for parameter-efficient fine-tuning
Pengjie Ren, Chengshun Shi, Shiguang Wu, Mengqi Zhang, Zhaochun Ren, Maarten de Rijke, Zhumin Chen, and Jiahuan Pei. 2024 · 2024
Later among the works it cites.
Multimodal instruction tuning with conditional mixture of LoRA
Ying Shen, Zhiyang Xu, Qifan Wang, Yu Cheng, Wenpeng Yin, and Lifu Huang. 2024 · 2024
Later among the works it cites.
ResLoRA: Identity residual mapping in low-rank adaption
Shuhua Shi, Shaohan Huang, Minghui Song, Zhoujun Li, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2024 · 2024
Later among the works it cites.
A simple and effective pruning approach for large language models
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter. 2024 · 2024
Later among the works it cites.
HydraloRA: An asymmetric loRA architecture for efficient fine-tuning
Chunlin Tian, Zhan Shi, Zhijiang Guo, Li Li, and Cheng zhong Xu. 2024 · 2024
Later among the works it cites.
Advancing parameter efficiency in fine-tuning via representation editing
Muling Wu, Wenhao Liu, Xiaohua Wang, Tianlong Li, Changze Lv, Zixuan Ling, Jianhao Zhu, Cenyuan Zhang, Xiaoqing Zheng, and Xuanjing Huang. 2024 · 2024
Later among the works it cites.
QA-loRA: Quantization-aware low-rank adaptation of large language models
Yuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen, Heng Chang, Hengheng Zhang, Zhengsu Chen, XIAOPENG ZHANG, and Qi Tian. 2024 · 2024
Later among the works it cites.
Functional faithfulness in the wild: Circuit discovery with differentiable computation graph pruning
Lei Yu, Jingcheng Niu, Zining Zhu, and Gerald Penn. 2024 · 2024
Later among the works it cites.
MiLoRA: Efficient mixture of low-rank adaptation for large language models fine-tuning
Jingfan Zhang, Yi Zhao, Dan Chen, Xing Tian, Huanran Zheng, and Wei Zhu. 2024 · 2024
Later among the works it cites.
t 2 t^{2} of thoughts: Temperature tree elicits reasoning in large language models
Chengkun Cai, Xu Zhao, Yucheng Du, Haoliang Liu, and Lei Li. 2025 · 2025
Closest in time.
Classic4Children: Adapting Chinese literary classics for children with large language model
Jiali Chen, Xusen Hei, Yuqi Xue, Zihan Wu, Jiayuan Xie, and Yi Cai. 2025 · 2025
Closest in time.
Everything is editable: Extend knowledge editing to unstructured data in large language models
Jingcheng Deng, Zihao Wei, Liang Pang, Hanxing Ding, Huawei Shen, and Xueqi Cheng. 2025 · 2025
Closest in time.
SLAM: Towards efficient multilingual reasoning via selective language alignment
Yuchun Fan, Yongyu Mu, YiLin Wang, Lei Huang, Junhao Ruan, Bei Li, Tong Xiao, Shujian Huang, Xiaocheng Feng, and Jingbo Zhu. 2025 · 2025
Closest in time.
Geoedit: Geometric knowledge editing for large language models
Yujie Feng, Liming Zhan, Zexin Lu, Yongxin Xu, Xu Chu, Yasha Wang, Jiannong Cao, Philip S Yu, and Xiao-Ming Wu. 2025 · 2025
Closest in time.
Beyond decoder-only: Large language models can be good encoders for machine translation
Yingfeng Luo, Tong Zheng, Yongyu Mu, Bei Li, Qinghong Zhang, Yongqi Gao, Ziqiang Xu, Peinan Feng, Xiaoqian Liu, Tong Xiao, et al. 2025 · 2025
Closest in time.
Problem-solving logic guided curriculum in-context learning for llms complex reasoning
Xuetao Ma, Wenbin Jiang, and Hua Huang. 2025 · 2025
Closest in time.
Uncertainty and influence aware reward model refinement for reinforcement learning from human feedback
Zexu Sun, Yiju Guo, Yankai Lin, Xu Chen, Qi Qi, Xing Tang, xiuqiang He, and Ji-Rong Wen. 2025 · 2025
Closest in time.
Text-to-cad generation through infusing visual feedback in large language models
Ruiyu Wang, Yu Yuan, Shizhao Sun, and Jiang Bian. 2025 · 2025
Closest in time.
Redstar: Does scaling long-cot data unlock better slow-reasoning systems?
Haotian Xu, Xing Wu, Weinong Wang, Zhongzhi Li, Da Zheng, Boyuan Chen, Yi Hu, Shijia Kang, Jiaming Ji, Yingying Zhang, Zhijiang Guo, Yaodong Yang, Muhan Zhang, and Debing Zhang. 2025 · 2025
Closest in time.
A survey on test-time scaling in large language models: What, how, where, and how well?
Qiyuan Zhang, Fuyuan Lyu, Zexu Sun, Lei Wang, Weixu Zhang, Wenyue Hua, Haolun Wu, Zhihan Guo, Yufei Wang, Niklas Muennighoff, Irwin King, Xue Liu, and Chen Ma. 2025 · 2025
Closest in time.
Asymmetric conflict and synergy in post-training for llm-based multilingual machine translation
Tong Zheng, Yan Wen, Huiwen Bao, Junfeng Guo, and Heng Huang. 2025 · 2025
Closest in time.
Are NLP models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021 · 2094
Closest in time.