Fetching the paper…
Reading the bibliography…
Low-rank Adaptation (LoRA) has demonstrated remarkable capabilities for task specific fine-tuning.
Chaos in random neural networks
Haim Sompolinsky, Andrea Crisanti, and Hans-Jurgen Sommers · 1988
Earlier work this paper cites.
Stochastic gradient learning in neural networks
Léon Bottou et al · 1991
Earlier work this paper cites.
Stochastic neural networks
Eugene Wong · 1991
Earlier work this paper cites.
Feed forward neural networks with random weights
Wouter F Schmidt, Martin A Kraaijveld, Robert PW Duin, et al · 1992
Earlier work this paper cites.
Network information criterion-determining the number of hidden units for an artificial neural network model
Noboru Murata, Shuji Yoshizawa, and Shun-ichi Amari · 1994
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
Bayesian learning for neural networks
Radford M Neal · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma · 2013
Earlier work this paper cites.
The icap framework: Linking cognitive engagement to active learning outcomes
Michelene TH Chi and Ruth Wylie · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Earlier work this paper cites.
Weight uncertainty in neural network
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Durk P Kingma, Tim Salimans, and Max Welling · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
David Ha, Andrew Dai, and Quoc V Le · 2016
Earlier work this paper cites.
Fm-delta: Fault management packet compression
Tal Mizrahi, Yoram Revah, Yehonathan Refael Kalim, Elad Kapuza, and Yuval Cassuto · 2017
Earlier work this paper cites.
Contextual parameter generation for universal neural machine translation
Emmanouil Antonios Platanios, Mrinmaya Sachan, Graham Neubig, and Tom Mitchell · 2018
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Cited alongside, same era.
Dataset condensation with gradient matching
Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen · 2020
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Hyperdiffusion: Generating implicit neural fields with weight-space diffusion
Ziya Erkoç, Fangchang Ma, Qi Shan, Matthias Nießner, and Angela Dai · 2023
Later among the works it cites.
In-context learning creates task vectors
Roee Hendel, Mor Geva, and Amir Globerson · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
P-diff: Learning classifier with noisy labels based on probability difference distributions
Wei Hu, QiHao Zhao, Yangyu Huang, and Fan Zhang · 2021
Cited alongside, same era.
Metaicl: Learning to learn in context
Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Dataset condensation with differentiable siamese augmentation
Bo Zhao and Hakan Bilen · 2021
Cited alongside, same era.
P-diff+: Improving learning classifier with noisy labels by noisy negative learning loss
QiHao Zhao, Wei Hu, Yangyu Huang, and Fan Zhang · 2021
Cited alongside, same era.
Dataset condensation with contrastive signals
Saehyung Lee, Sanghyuk Chun, Sangwon Jung, Sangdoo Yun, and Sungroh Yoon · 2022
Cited alongside, same era.
Learning to learn with generative models of neural network checkpoints
William Peebles, Ilija Radosavovic, Tim Brooks, Alexei A Efros, and Jitendra Malik · 2022
Cited alongside, same era.
Later among the works it cites.
Small models are valuable plug-ins for large language models
Canwen Xu, Yichong Xu, Shuohang Wang, Yang Liu, Chenguang Zhu, and Julian McAuley · 2023
Later among the works it cites.
Dataset condensation with distribution matching
Bo Zhao and Hakan Bilen · 2023
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Multisize dataset condensation
Yang He, Lingao Xiao, Joey Tianyi Zhou, and Ivor Tsang · 2024
Later among the works it cites.
In-context lora for diffusion transformers
Lianghua Huang, Wei Wang, Zhi-Fan Wu, Yupeng Shi, Huanzhang Dou, Chen Liang, Yutong Feng, Yu Liu, and Jingren Zhou · 2024
Later among the works it cites.
Conditional lora parameter generation
Xiaolong Jin, Kai Wang, Dongwen Tang, Wangbo Zhao, Yukun Zhou, Junshu Tang, and Yang You · 2024
Later among the works it cites.
Dataset condensation for time series classification via dual domain matching
Zhanyu Liu, Ke Hao, Guanjie Zheng, and Yanwei Yu · 2024
Later among the works it cites.
Gwq: Gradient-aware weight quantization for large language models
Yihua Shao, Siyu Liang, Xiaolin Lin, Zijian Ling, Zixian Zhu, Minxi Yan, Haiyang Liu, Siyu Chen, Ziyang Yan, Yilan Meng, et al · 2024
Later among the works it cites.
Dataset condensation with latent quantile matching
Wei Wei, Tom De Schepper, and Kevin Mets · 2024
Later among the works it cites.
Florence-2: Advancing a unified representation for a variety of vision tasks
Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, and Lu Yuan · 2024
Later among the works it cites.
3dsceneeditor: Controllable 3d scene editing with gaussian splatting
Ziyang Yan, Lei Li, Yihua Shao, Siyu Chen, Wuzong Kai, Jenq-Neng Hwang, Hao Zhao, and Fabio Remondino · 2024
Later among the works it cites.