Fetching the paper…
Reading the bibliography…
Foundation models (FM) have demonstrated remarkable performance across a wide range of tasks (especially in the fields of natural language processing and computer vision), primarily attributed to their ability to comprehend instructions and access extensive, high-quality data.
Variational Information Distillation for Knowledge Transfer, April 2019
S. Ahn et al · 1904
Earlier work this paper cites.
Learning a synaptic learning rule
Y. Bengio et al · 1990
Earlier work this paper cites.
Experiments on multistrategy learning by meta-learning
P. K. Chan and S. J. Stolfo · 1993
Earlier work this paper cites.
On the search for new learning rules for anns
S. Bengio et al · 1995
Earlier work this paper cites.
Simple principles of metalearning
J. Schmidhuber et al · 1996
Earlier work this paper cites.
Ensemble methods in machine learning
T. G. Dietterich · 2000
Earlier work this paper cites.
A perspective view and survey of meta-learning
R. Vilalta and Y. Drissi · 2002
Earlier work this paper cites.
Model as a service (maas)
D. Roman et al · 2009
Earlier work this paper cites.
Kullback-leibler divergence
J. M. Joyce · 2011
Earlier work this paper cites.
Learning to learn
S. Thrun and L. Pratt · 2012
Earlier work this paper cites.
Over-fitting and model tuning
M. Kuhn et al · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow et al · 2014
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network, March 2015
G. Hinton et al · 2015
Earlier work this paper cites.
FitNets: Hints for Thin Deep Nets, March 2015
A. Romero et al · 2015
Earlier work this paper cites.
Inceptionism: Going deeper into neural networks
A. Mordvintsev et al · 2015
Earlier work this paper cites.
Convolutional neural networks for medical image analysis: Full training or fine tuning?
N. Tajbakhsh et al · 2016
Earlier work this paper cites.
Improving neural language models with a continuous cache
E. Grave et al · 2016
Earlier work this paper cites.
A Chinese Question Answering Approach Integrating Count-Based and Embedding-Based Features
B. Wang et al · 2016
Earlier work this paper cites.
Optimization as a model for few-shot learning
S. Ravi and H. Larochelle · 2016
Earlier work this paper cites.
Fine-tuning convolutional neural networks for biomedical image analysis: Actively and incrementally
Z. Zhou et al · 2017
Earlier work this paper cites.
Data-free knowledge distillation for deep neural networks
R. G. Lopes et al · 2017
Earlier work this paper cites.
Deep forest: Towards an alternative to deep neural networks
Z. Zhou and J. Feng · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn et al · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford et al · 2018
Earlier work this paper cites.
Learning Student Networks via Feature Embedding, December 2018
H. Chen et al · 2018
Earlier work this paper cites.
Ensemble learning: A survey
O. Sagi and L. Rokach · 2018
Earlier work this paper cites.
Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs, October 2018
T. Garipov et al · 2018
Earlier work this paper cites.
On first-order meta-learning algorithms
A. Nichol et al · 2018
Earlier work this paper cites.
Learning to compare: Relation network for few-shot learning
F. Sung et al · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
N. Houlsby et al · 2019
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
U. Khandelwal et al · 2019
Earlier work this paper cites.
LIT: Learned Intermediate Representation Training for Model Compression
A. Koratana et al · 2019
Earlier work this paper cites.
Language models as knowledge bases?
F. Petroni et al · 2019
Earlier work this paper cites.
Dream distillation: A data-independent model compression framework
K. Bhardwaj et al · 2019
Earlier work this paper cites.
Zero-shot knowledge distillation in deep networks
G. K. Nayak et al · 2019
Earlier work this paper cites.
Deep leakage from gradients
L. Zhu et al · 2019
Earlier work this paper cites.
Data-free learning of student networks
H. Chen et al · 2019
Earlier work this paper cites.
Knowledge extraction with no observable data
J. Yoo et al · 2019
Earlier work this paper cites.
Zero-shot knowledge transfer via adversarial belief matching
P. Micaelli and A. J. Storkey · 2019
Earlier work this paper cites.
Data-free adversarial distillation
G. Fang et al · 2019
Earlier work this paper cites.
Ensemble Approach for Natural Language Question Answering Problem
A. Aniol et al · 2019
Earlier work this paper cites.
Averaging Weights Leads to Wider Optima and Better Generalization, February 2019
P. Izmailov et al · 2019
Earlier work this paper cites.
Learning to learn how to learn: Self-adaptive visual navigation using meta-learning
M. Wortsman et al · 2019
Earlier work this paper cites.
Meta-transfer learning for few-shot learning
Q. Sun et al · 2019
Earlier work this paper cites.
Query-efficient meta attack to deep neural networks
J. Du et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown et al · 2020
Earlier work this paper cites.
How can we know what language models know?
Z. Jiang et al · 2020
Earlier work this paper cites.
Big self-supervised models are strong semi-supervised learners
T. Chen et al · 2020
Earlier work this paper cites.
Better fine-tuning by reducing representational collapse
A. Aghajanyan et al · 2020
Earlier work this paper cites.
Deep learning on computational-resource-limited platforms: a survey
C. Chen et al · 2020
Earlier work this paper cites.
Remind your neural network to prevent catastrophic forgetting
T. L. Hayes et al · 2020
Earlier work this paper cites.
Adapterfusion: Non-destructive task composition for transfer learning
J. Pfeiffer et al · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
K. Guu et al · 2020
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering
G. Izacard and E. Grave · 2020
Earlier work this paper cites.
Distilling knowledge from reader to retriever for question answering
G. Izacard and E. Grave · 2020
Earlier work this paper cites.
Distilling Cross-Task Knowledge via Relationship Matching
H.-J. Ye et al · 2020
Earlier work this paper cites.
Dreaming to distill: Data-free knowledge transfer via deepinversion
H. Yin et al · 2020
Earlier work this paper cites.
Inverting gradients-how easy is it to break privacy in federated learning?
J. Geiping et al · 2020
Earlier work this paper cites.
Data-free knowledge amalgamation via group-stack dual-gan
J. Ye et al · 2020
Earlier work this paper cites.
Degan: Data-enriching gan for retrieving representative samples from a trained classifier
S. Addepalli et al · 2020
Earlier work this paper cites.
BatchEnsemble: An Alternative Approach to Efficient Ensemble and Lifelong Learning, February 2020
Y. Wen et al · 2020
Earlier work this paper cites.
Linear Mode Connectivity and the Lottery Ticket Hypothesis, July 2020
J. Frankle et al · 2020
Earlier work this paper cites.
Towards Efficient Front-End Visual Sensing for Digital Retina: A Model-Centric Paradigm
Y. Lou et al · 2020
Earlier work this paper cites.
Online fast adaptation and knowledge accumulation (osaka): a new approach to continual learning
M. Caccia et al · 2020
Earlier work this paper cites.
Meta-transfer learning through hard tasks
Q. Sun et al · 2020
Earlier work this paper cites.
Codebert: A pre-trained model for programming and natural languages
Z. Feng et al · 2020
Earlier work this paper cites.
Modeling and optimization trade-off in meta-learning
K. Gao and O. Sener · 2020
Earlier work this paper cites.
A. Sinitsin et al · 2020
Earlier work this paper cites.
Modifying memories in transformer models, 2020
C. Zhu et al · 2020
Earlier work this paper cites.
Knowledge distillation: A survey
J. Gou et al · 2021
Earlier work this paper cites.
Extracting training data from large language models
N. Carlini et al · 2021
Earlier work this paper cites.
Meta-learning in neural networks: A survey
T. Hospedales et al · 2021
Earlier work this paper cites.
E. Mitchell et al · 2021
Earlier work this paper cites.
Lightweight adapter tuning for multilingual speech translation
H. Le et al · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation, 2021
X. L. Li and P. Liang · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
B. Lester et al · 2021
Cited alongside, same era.
Recent advances in natural language processing via large pre-trained language models: A survey
B. Min et al · 2021
Cited alongside, same era.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
E. B. Zaken et al · 2021
Cited alongside, same era.
On the effectiveness of adapter-based tuning for pretrained language model adaptation
R. He et al · 2021
Cited alongside, same era.
Compacter: Efficient low-rank hypercomplex adapter layers
R. Karimi Mahabadi et al · 2021
Cited alongside, same era.
Calibrating factual knowledge in pretrained language models
Q. Dong et al · 2022
Later among the works it cites.
Fixing model bugs with natural language patches
S. Murty et al · 2022
Later among the works it cites.
Mass-editing memory in a transformer
K. Meng et al · 2022
Later among the works it cites.
Editing models with task arithmetic
G. Ilharco et al · 2022
Later among the works it cites.
Locating and editing factual associations in gpt
K. Meng et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The power of scale for parameter-efficient prompt tuning
B. Lester et al · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
X. L. Li and P. Liang · 2021
Cited alongside, same era.
Learning how to ask: Querying LMs with mixtures of soft prompts
G. Qin and J. Eisner · 2021
Cited alongside, same era.
X. Liu et al · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
J. Wei et al · 2021
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
V. Sanh et al · 2021
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks, 2021
P. Lewis et al · 2021
Cited alongside, same era.
V. Raunak and A. Menezes · 2022
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron et al · 2023
Closest in time.
Why johnny can’t prompt: how non-ai experts try (and fail) to design llm prompts
J. Zamfirescu-Pereira et al · 2023
Closest in time.
Large language models (llm) and chatgpt: what will the impact on nuclear medicine be?
I. L. Alberts et al · 2023
Closest in time.
Text-to-audio generation using instruction-tuned llm and latent diffusion model
D. Ghosal et al · 2023
Closest in time.
Summary of chatgpt/gpt-4 research and perspective towards the future of large language models
Y. Liu et al · 2023
Closest in time.
Model robustness meets data privacy: Adversarial robustness distillation without original data
Y. Wang et al · 2023
Closest in time.
D. Misra et al · 2023
Closest in time.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
D. Jiang et al · 2023
Closest in time.
Multi-head adapter routing for cross-task generalization, 2023
L. Caccia et al · 2023
Closest in time.
Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models
Z. Hu et al · 2023
Closest in time.
Black box adversarial prompting for foundation models, 2023
N. Maus et al · 2023
Closest in time.
Progressive prompts: Continual learning for language models
A. Razdaibiedina et al · 2023
Closest in time.
Gradient-regulated meta-prompt learning for generalizable vision-language models, 2023
J. Li et al · 2023
Closest in time.
Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning, 2023
Z. Xu et al · 2023
Closest in time.
H. Liu et al · 2023
Closest in time.
Gpt4roi: Instruction tuning large language model on region-of-interest
S. Zhang et al · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
P. Liu et al · 2023
Closest in time.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
N. Ruiz et al · 2023
Closest in time.
On the effectiveness of parameter-efficient fine-tuning
Z. Fu et al · 2023
Closest in time.
Parameter-efficient fine-tuning of large-scale pre-trained language models
N. Ding et al · 2023
Closest in time.
Exploring efficient-tuning methods in self-supervised speech models
Z.-C. Chen et al · 2023
Closest in time.
Using adapters to overcome catastrophic forgetting in end-to-end automatic speech recognition
S. Vander Eeckt and H. Van Hamme · 2023
Closest in time.
Rethinking efficient tuning methods from a unified perspective
Z. Jiang et al · 2023
Closest in time.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Y. Wen et al · 2023
Closest in time.
Offsite-tuning: Transfer learning without full model
G. Xiao et al · 2023
Closest in time.
The flan collection: Designing data and methods for effective instruction tuning
S. Longpre et al · 2023
Closest in time.
B. Peng et al · 2023
Closest in time.
Exploring the benefits of training expert language models over instruction tuning, 2023
J. Jang et al · 2023
Closest in time.
Otter: A multi-modal model with in-context instruction tuning
B. Li et al · 2023
Closest in time.
Shall we pretrain autoregressive language models with retrieval? a comprehensive study
B. Wang et al · 2023
Closest in time.
Replug: Retrieval-augmented black-box language models, 2023
W. Shi et al · 2023
Closest in time.
Retrieval-augmented multimodal language modeling
M. Yasunaga et al · 2023
Closest in time.
Deep Classifier Mimicry without Data Access, June 2023
S. Braun et al · 2023
Closest in time.
Learning to Learn from APIs: Black-Box Data-Free Meta-Learning
Z. Hu et al · 2023
Closest in time.
IS SYNTHETIC DATA FROM GENERATIVE MODELS READY FOR IMAGE RECOGNITION?
R. He et al · 2023
Closest in time.
Learning to retain while acquiring: Combating distribution-shift in adversarial data-free knowledge distillation
G. Patel et al · 2023
Closest in time.
Tangent Model Composition for Ensembling and Continual Fine-tuning, July 2023
T. Y. Liu and S. Soatto · 2023
Closest in time.
W. Li et al · 2023
Closest in time.
AdapterSoup: Weight Averaging to Improve Generalization of Pretrained Language Models, March 2023
A. Chronopoulou et al · 2023
Closest in time.
Understanding the Effectiveness of Early Weight Averaging for Training Large Language Models, June 2023
S. Sanyal et al · 2023
Closest in time.
Communication-Efficient Learning of Deep Networks from Decentralized Data, January 2023
H. B. McMahan et al · 2023
Closest in time.
PopulAtion Parameter Averaging (PAPA), May 2023
A. Jolicoeur-Martineau et al · 2023
Closest in time.
Editing models with task arithmetic
G. Ilharco et al · 2023
Closest in time.
Git Re-Basin: Merging Models modulo Permutation Symmetries, March 2023
S. K. Ainsworth et al · 2023
Closest in time.
ZipIt! Merging Models from Different Tasks without Training, May 2023
George Stoica et al · 2023
Closest in time.
Model Fusion via Optimal Transport, May 2023
S. P. Singh and M. Jaggi · 2023
Closest in time.
Dataless Knowledge Fusion by Merging Weights of Language Models, April 2023
X. Jin et al · 2023
Closest in time.
C. Wu et al · 2023
Closest in time.
Boosting the Adversarial Transferability of Surrogate Models with Dark Knowledge, September 2023
D. Yang et al · 2023
Closest in time.
NTK-approximating MLP Fusion for Efficient Language Model Fine-tuning, July 2023
T. Wei et al · 2023
Closest in time.
Tangent Model Composition for Ensembling and Continual Fine-tuning, July 2023
T. Y. Liu and S. Soatto · 2023
Closest in time.
Federated Learning of Shareable Bases for Personalization-Friendly Image Classification, April 2023
H.-Y. Chen et al · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Learning to learn from apis: Black-box data-free meta-learning
Z. Hu et al · 2023
Closest in time.
Architecture, dataset and model-scale agnostic data-free meta-learning
Z. Hu et al · 2023
Closest in time.
Film: How can few-shot image classification benefit from pre-trained language models?
Z. Jiang et al · 2023
Closest in time.
Training meta-surrogate model for transferable adversarial attack
Y. Qin et al · 2023
Closest in time.
D-dae: Defense-penetrating model extraction attacks
Y. Chen et al · 2023
Closest in time.
Meta-learning-based optimal control for soft robotic manipulators to interact with unknown environments
Z. Tang et al · 2023
Closest in time.
S. Watanabe et al · 2023
Closest in time.
An overview on meta-learning approaches for few-shot weakly-supervised segmentation
P. H. T. Gama et al · 2023
Closest in time.
Editing large language models: Problems, methods, and opportunities
Y. Yao et al · 2023
Closest in time.
Memory-assisted prompt editing to improve gpt-3 after deployment, 2023
A. Madaan et al · 2023
Closest in time.
Transformer-patcher: One mistake worth one neuron
Z. Huang et al · 2023
Closest in time.
Can lms learn new entities from descriptions? challenges in propagating injected knowledge
Y. Onoe et al · 2023
Closest in time.
Prompt-based editing for text style transfer
G. Luo et al · 2023
Closest in time.
Conditional text image generation with diffusion models
Y. Zhu et al · 2023
Closest in time.
Crawling the internal knowledge-base of language models, 2023
R. Cohen et al · 2023
Closest in time.
The life cycle of knowledge in big language models: A survey
B. Cao et al · 2023
Closest in time.