Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have become increasingly capable, but their development often requires substantial computational resources.
Visual Question Answering Dataset for Bilingual Image Understanding: A Study of Cross-Lingual Transfer Using Attention Maps. In
Nobuyuki Shimizu, Na Rong, and Takashi Miyazaki. 2018 · 1928
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Jürgen Schmidhuber. 1992 · 1992
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Sepp Hochreiter and Jürgen Schmidhuber. 1994 · 1994
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
The cathedral and the bazaar
Eric Raymond. 1999 · 1999
Earlier work this paper cites.
A fast and elitist multiobjective genetic algorithm: NSGA-II
Kalyanmoy Deb, Amrit Pratap, Sameer Agarwal, and TAMT Meyarivan. 2002 · 2002
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
Kenneth O Stanley and Risto Miikkulainen. 2002 · 2002
Earlier work this paper cites.
The CMA evolution strategy: a comparing review
Nikolaus Hansen. 2006 · 2006
Earlier work this paper cites.
David Ha, Andrew Dai, and Quoc V Le. 2016 · 2016
Earlier work this paper cites.
FastText.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. 2016b · 2016
Earlier work this paper cites.
Bag of Tricks for Efficient Text Classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2016a · 2016
Earlier work this paper cites.
Tom White. 2016 · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le. 2016 · 2016
Earlier work this paper cites.
Gintare Karolina Dziugaite and Daniel M Roy. 2017 · 2017
Earlier work this paper cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. In
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2017 · 2017
Earlier work this paper cites.
Optuna: A Next-generation Hyperparameter Optimization Framework. In
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019 · 2019
Earlier work this paper cites.
Weight agnostic neural networks
Adam Gaier and David Ha. 2019 · 2019
Earlier work this paper cites.
Regularized evolution for image classifier architecture search. In
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le. 2019 · 2019
Earlier work this paper cites.
The evolved transformer. In
David So, Quoc Le, and Chen Liang. 2019 · 2019
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Relative Flatness and Generalization. In
Henning Petzka, Michael Kamp, Linara Adilova, Cristian Sminchisescu, and Mario Boley. 2021 · 2021
Cited alongside, same era.
A call to build models like we build open-source software
Colin Raffel. 2021 · 2021
Cited alongside, same era.
Stable Diffusion WebUI
AUTOMATIC1111. 2022 · 2022
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg. 2022 · 2022
Cited alongside, same era.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2022 · 2022
Cited alongside, same era.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 · 2023
Later among the works it cites.
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 2023 · 2023
Later among the works it cites.
Building machine learning models like open source software
Colin Raffel. 2023 · 2023
Later among the works it cites.
Language models are multilingual chain-of-thought reasoners. In
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2023 · 2023
Later among the works it cites.
Japanese Stable VLM
Makoto Shing and Takuya Akiba. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jean Kaddour, Linqing Liu, Ricardo Silva, and Matt J Kusner. 2022 · 2022
Cited alongside, same era.
Merging models with fisher-weighted averaging
Michael S Matena and Colin A Raffel. 2022 · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models. In
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Cited alongside, same era.
EvoJAX: Hardware-Accelerated Neuroevolution
Yujin Tang, Yingtao Tian, and David Ha. 2022 · 2022
Cited alongside, same era.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al · 2022
Cited alongside, same era.
GPT-4V(ision) System Card
Open AI. 2023 · 2023
Cited alongside, same era.
An empirical study of multimodal model merging
Yi-Lin Sung, Linjie Li, Kevin Lin, Zhe Gan, Mohit Bansal, and Lijuan Wang. 2023 · 2023
Later among the works it cites.
Prateek Yadav, Leshem Choshen, Colin Raffel, and Mohit Bansal. 2023a · 2023
Later among the works it cites.
TIES-Merging: Resolving Interference When Merging Models. In
Prateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel, and Mohit Bansal. 2023b · 2023
Later among the works it cites.
Model Merging by Uncertainty-Based Gradient Matching. In
Nico Daheim, Thomas Möllenhoff, Edoardo Ponti, Iryna Gurevych, and Mohammad Emtiyaz Khan. 2024 · 2024
Closest in time.
mergekit
Charles O. Goddard. 2024 · 2024
Closest in time.
Automerger Experiment
Maxime Labonne. 2024a · 2024
Closest in time.
Merge Large Language Models with mergekit
Maxime Labonne. 2024b · 2024
Closest in time.
SDXL-Lightning: Progressive Adversarial Diffusion Distillation
Shanchuan Lin, Anran Wang, and Xiao Yang. 2024 · 2024
Closest in time.
LLaVA-NeXT: Improved reasoning, OCR, and world knowledge
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024 · 2024
Closest in time.
Interpreting GPT: The Logit Lens
nostalgebraist. 2021 · 2024
Closest in time.
LM Benchmark
rinna. 2024 · 2024
Closest in time.
Transformer Layers as Painters
Qi Sun, Marc Pickett, Aakash Kumar Nain, and Llion Jones. 2024 · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al · 2024
Closest in time.
Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li. 2024 · 2024
Closest in time.