Fetching the paper…
Reading the bibliography…
Recent advances in foundation models have emphasized the need to align pre-trained models with specialized domains using small, curated datasets.
A Simple General Approach to Inference About the Tail of a Distribution
Bruce M. Hill · 1975
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation, 2015
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Chemprot-3.0: a global chemical biology diseases mapping
Jens Kringelum, Sonny Kim Kjaerulff, Søren Brunak, Ole Lund, Tudor I Oprea, and Olivier Taboureau · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Pubmed 200k rct: a dataset for sequential sentence classification in medical abstracts
Franck Dernoncourt and Ji Young Lee · 2017
Earlier work this paper cites.
Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost, 2018
Noam Shazeer and Mitchell Stern · 2018
Earlier work this paper cites.
Structural scaffolds for citation intent classification in scientific publications
Arman Cohan, Waleed Ammar, Madeleine Van Zuylen, and Field Cady · 2019
Earlier work this paper cites.
Semeval-2019 task 4: Hyperpartisan news detection
Johannes Kiesel, Maria Mestre, Rishabh Shukla, Emmanuel Vincent, Payam Adineh, David Corney, Benno Stein, and Martin Potthast · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations
Maziar Raissi, Paris Perdikaris, and George E Karniadakis · 2019
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding, 2019
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Fourier neural operator for parametric partial differential equations
Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar · 2020
Earlier work this paper cites.
Heavy-tailed universality predicts trends in test accuracies for very large pre-trained deep neural networks
Charles H Martin and Michael W Mahoney · 2020
Earlier work this paper cites.
On 1/n neural representation and robustness
Josue Nassar, Piotr Sokol, SueYeon Chung, Kenneth D Harris, and Il Memming Park · 2020
Earlier work this paper cites.
Hausdorff dimension, heavy tails, and generalization in neural networks
Umut Simsekli, Ozan Sener, George Deligiannidis, and Murat A Erdogdu · 2020
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems, 2020
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2020
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing, 2020
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Earlier work this paper cites.
Heavy tails in sgd and compressibility of overparametrized neural networks
Melih Barsbey, Milad Sefidgaran, Murat A Erdogdu, Gael Richard, and Umut Simsekli · 2021
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization, 2021
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Cited alongside, same era.
The heavy-tail phenomenon in sgd
Mert Gurbuzbalaban, Umut Simsekli, and Lingjiong Zhu · 2021
Cited alongside, same era.
Multiplicative noise and heavy tails in stochastic optimization
Liam Hodgkinson and Michael Mahoney · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Physics-informed machine learning
George Em Karniadakis, Ioannis G Kevrekidis, Lu Lu, Paris Perdikaris, Sifan Wang, and Liu Yang · 2021
Astroclip: Cross-modal pre-training for astronomical foundation models
Francois Lanusse, Liam Parker, Siavash Golkar, Miles Cranmer, Alberto Bietti, Michael Eickenberg, Geraud Krawezik, Michael McCabe, Ruben Ohana, Mariel Pettee, et al · 2023
Later among the works it cites.
Multiple physics pretraining for physical surrogate models
Michael McCabe, Bruno Régaldo-Saint Blancard, Liam Holden Parker, Ruben Ohana, Miles Cranmer, Alberto Bietti, Michael Eickenberg, Siavash Golkar, Geraud Krawezik, Francois Lanusse, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cutting down on prompts and parameters: Simple few-shot learning with language models
Robert L Logan IV, Ivana Balažević, Eric Wallace, Fabio Petroni, Sameer Singh, and Sebastian Riedel · 2021
Cited alongside, same era.
Learning nonlinear operators via deeponet based on the universal approximation theorem of operators
Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis · 2021
Cited alongside, same era.
Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning
Charles H Martin and Michael W Mahoney · 2021
Cited alongside, same era.
Predicting trends in the quality of state-of-the-art neural networks without access to training or testing data
Charles H Martin, Tongsu Peng, and Michael W Mahoney · 2021
Cited alongside, same era.
Revisiting few-sample {bert} fine-tuning
Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q Weinberger, and Yoav Artzi · 2021
Cited alongside, same era.
$\alpha$-req : Assessing {\bf Re}presentation {\bf Q}uality in self-supervised learning by measuring eigenspectrum decay
Kumar Krishna Agrawal, Arnab Kumar Mondal, Arna Ghosh, and Blake Aaron Richards · 2022
Cited alongside, same era.
Scientific discovery in the age of artificial intelligence
Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, Ziming Liu, Payal Chandak, Shengchao Liu, Peter Van Katwyk, Andreea Deac, et al · 2023
Later among the works it cites.
Towards generalist foundation model for radiology
Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie · 2023
Later among the works it cites.
Test accuracy vs. generalization gap: Model selection in nlp without accessing training or testing data
Yaoqing Yang, Ryan Theisen, Liam Hodgkinson, Joseph E Gonzalez, Kannan Ramchandran, Charles H Martin, and Michael W Mahoney · 2023
Later among the works it cites.
Lima: Less is more for alignment, 2023
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy · 2023
Later among the works it cites.
Data-efficient operator learning via unsupervised pretraining and in-context learning
Wuyang Chen, Jialin Song, Pu Ren, Shashank Subramanian, Dmitriy Morozov, and Michael W Mahoney · 2024
Closest in time.
Dpot: Auto-regressive denoising operator transformer for large-scale pde pre-training
Zhongkai Hao, Chang Su, Songming Liu, Julius Berner, Chengyang Ying, Hang Su, Anima Anandkumar, Jian Song, and Jun Zhu · 2024
Closest in time.
Crafting heavy-tails in weight matrix spectrum without gradient noise, 2024
Vignesh Kothapalli, Tianyu Pang, Shenyang Deng, Zongmin Liu, and Yaoqing Yang · 2024
Closest in time.
Owlore: Outlier-weighed layerwise sampled low-rank projection for memory-efficient llm fine-tuning
Pengxiang Li, Lu Yin, Xiaowei Gao, and Shiwei Liu · 2024
Closest in time.
Alphapruning: Using heavy-tailed self regularization theory for improved layer-wise pruning of large language models
Haiquan Lu, Yefan Zhou, Shiwei Liu, Zhangyang Wang, Michael W. Mahoney, and Yaoqing Yang · 2024
Closest in time.
Alphaexpert: Assigning lora experts based on layer training quality
Peijun Qing, Chongyang Gao, Yefan Zhou, Xingjian Diao, Yaoqing Yang, and Vosoughi Soroush · 2024
Closest in time.
Convolutional neural operators for robust and accurate learning of pdes
Bogdan Raonic, Roberto Molinaro, Tim De Ryck, Tobias Rohner, Francesca Bartolucci, Rima Alaifari, Siddhartha Mishra, and Emmanuel de Bézenac · 2024
Closest in time.
Towards foundation models for scientific machine learning: Characterizing scaling and transfer behavior
Shashank Subramanian, Peter Harrington, Kurt Keutzer, Wahid Bhimji, Dmitriy Morozov, Michael W Mahoney, and Amir Gholami · 2024
Closest in time.
On the overlooked structure of stochastic gradients
Zeke Xie, Qian-Yuan Tang, Mingming Sun, and Ping Li · 2024
Closest in time.
Pde generalization of in-context operator networks: A study on 1d scalar nonlinear conservation laws
Liu Yang and Stanley J Osher · 2024
Closest in time.
Pdeformer: Towards a foundation model for one-dimensional partial differential equations
Zhanhong Ye, Xiang Huang, Leheng Chen, Hongsheng Liu, Zidong Wang, and Bin Dong · 2024
Closest in time.
Temperature balancing, layer-wise weight analysis, and neural network training
Yefan Zhou, Tianyu Pang, Keqin Liu, Michael W Mahoney, Yaoqing Yang, et al · 2024
Closest in time.