Fetching the paper…
Reading the bibliography…
Instruction tuning has unlocked powerful capabilities in large language models (LLMs), effectively using combined datasets to develop generalpurpose chatbots.
The influence curve and its role in robust estimation
Hampel, F. R · 1974
Earlier work this paper cites.
Extensions of lipschitz mappings into hilbert space
Johnson, W. B. and Lindenstrauss, J · 1984
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Robertson, S., Zaragoza, H., et al · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
A latent variable model approach to pmi-based word embeddings
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Earlier work this paper cites.
Coresets and sketches
Phillips, J. M · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Sener, O. and Savarese, S · 2018
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
Toneva, M., Sordoni, A., des Combes, R. T., Trischler, A., Bengio, Y., and Gordon, G. J · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., and Wang, R · 2019
Earlier work this paper cites.
Selection via proxy: Efficient data selection for deep learning
Coleman, C., Yeh, C., Mussmann, S., Mirzasoleiman, B., Bailis, P., Liang, P., Leskovec, J., and Zaharia, M · 2019
Earlier work this paper cites.
Learning from less data: A unified data subset selection and active learning framework for computer vision
Kaushal, V., Iyer, R., Kothawade, S., Mahadev, R., Doctor, K., and Ramakrishnan, G · 2019
Earlier work this paper cites.
Influence functions in deep learning are fragile
Basu, S., Pope, P., and Feizi, S · 2020
Earlier work this paper cites.
TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages
Clark, J. H., Choi, E., Collins, M., Garrette, D., Kwiatkowski, T., Nikolaev, V., and Palomaki, J · 2020
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
Feldman, V. and Zhang, C · 2020
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A · 2020
Earlier work this paper cites.
Evaluation of similarity-based explanations
Hanawa, K., Yokoi, S., Hara, S., and Inui, K · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Hu, J., Ruder, S., Siddhant, A., Neubig, G., Firat, O., and Johnson, M · 2020
Earlier work this paper cites.
Coresets for data-efficient training of machine learning models
Mirzasoleiman, B., Bilmes, J., and Leskovec, J · 2020
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
Pruthi, G., Liu, F., Kale, S., and Sundararajan, M · 2020
Earlier work this paper cites.
Optimizing data usage via differentiable rewards
Wang, X., Pham, H., Michel, P., Anastasopoulos, A., Carbonell, J., and Neubig, G · 2020
Earlier work this paper cites.
Predicting performance for natural language processing tasks
Xia, M., Anastasopoulos, A., Xu, R., Yang, Y., and Neubig, G · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Trivial or impossible—dichotomous data difficulty masks model differences (on imagenet and beyond)
Meding, K., Buschoff, L. M. S., Geirhos, R., and Wichmann, F. A · 2021
Cited alongside, same era.
Dataset meta-learning from kernel ridge-regression
Nguyen, T., Chen, Z., and Lee, J · 2021
Cited alongside, same era.
Deep learning on a data diet: Finding important examples early in training
Paul, M., Ganguli, S., and Dziugaite, G. K · 2021
Cited alongside, same era.
Enhancing chat language models by scaling high-quality instructional conversations
Ding, N., Chen, Y., Xu, B., Qin, Y., Zheng, Z., Hu, S., Liu, Z., Sun, M., and Zhou, B · 2023
Later among the works it cites.
Mods: Model-oriented data selection for instruction tuning
Du, Q., Zong, C., and Zhang, J · 2023
Later among the works it cites.
An important next step on our ai journey, 2023
Google · 2023
Later among the works it cites.
Studying large language model generalization with influence functions, 2023
Grosse, R., Bae, J., Anil, C., Elhage, N., Tamkin, A., Tajdini, A., Steiner, B., Li, D., Durmus, E., Perez, E., Hubinger, E., Lukošiūtė, K., Nguyen, K., Joseph, N., McCandlish, S., Kaplan, J., and Bowman, S. R · 2023
Later among the works it cites.
Simfluence: Modeling the influence of individual training examples by simulating training runs, 2023
Guu, K., Webson, A., Pavlick, E., Dixon, L., Tenney, I., and Bolukbasi, T · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Revisiting methods for finding influential examples
Søgaard, A. et al · 2021
Cited alongside, same era.
Scale efficiently: Insights from pretraining and finetuning transformers
Tay, Y., Dehghani, M., Rao, J., Fedus, W., Abnar, S., Chung, H. W., Narang, S., Yogatama, D., Vaswani, A., and Metzler, D · 2021
Cited alongside, same era.
Tensor programs iv: Feature learning in infinite-width neural networks
Yang, G. and Hu, E. J · 2021
Cited alongside, same era.
If influence functions are the answer, then what is the question?
Bae, J., Ng, N. H., Lo, A., Ghassemi, M., and Grosse, R. B · 2022
Cited alongside, same era.
Datamodels: Predicting predictions from training data
Ilyas, A., Park, S. M., Engstrom, L., Leclerc, G., and Madry, A · 2022
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Cited alongside, same era.
Post-hoc interpretability for neural nlp: A survey
Madsen, A., Reddy, S., and Chandar, S · 2022
Cited alongside, same era.
Later among the works it cites.
In-context alignment: Chat with vanilla language models before fine-tuning
Han, X · 2023
Later among the works it cites.
Understanding in-context learning via supportive pretraining data
Han, X., Simig, D., Mihaylov, T., Tsvetkov, Y., Celikyilmaz, A., and Wang, T · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Later among the works it cites.
OpenAssistant conversations–democratizing large language model alignment
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z.-R., Stevens, K., Barhoum, A., Duc, N. M., Stanley, O., Nagyfi, R., et al · 2023
Later among the works it cites.
One shot learning as instruction data prospector for large language models
Li, Y., Hui, B., Xia, X., Yang, J., Yang, M., Zhang, L., Si, S., Liu, J., Liu, T., Huang, F., et al · 2023
Later among the works it cites.
Liu, W., Zeng, W., He, K., Jiang, Y., and He, J · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Later among the works it cites.
A kernel-based view of language model fine-tuning
Malladi, S., Wettig, A., Yu, D., Chen, D., and Arora, S · 2023
Later among the works it cites.
Orca: Progressive learning from complex explanation traces of gpt-4
Mukherjee, S., Mitra, A., Jawahar, G., Agarwal, S., Palangi, H., and Awadallah, A · 2023
Later among the works it cites.
OpenAI: GPT-4, 2023
OpenAI · 2023
Later among the works it cites.
Trak: Attributing model behavior at scale
Park, S. M., Georgiev, K., Ilyas, A., Leclerc, G., and Madry, A · 2023
Later among the works it cites.
Understanding influence functions and datamodels via harmonic analysis
Saunshi, N., Gupta, A., Braverman, M., and Arora, S · 2023
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q., Chi, E., Zhou, D., et al · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Moderate coreset: A universal method of data selection for real-world data-efficient deep learning
Xia, X., Liu, J., Yu, J., Shen, X., Han, B., and Liu, T · 2023
Later among the works it cites.
WizardLM: Empowering large language models to follow complex instructions
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., and Jiang, D · 2023
Later among the works it cites.
LIMA: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2023
Later among the works it cites.
Dsdm: Model-aware dataset selection with datamodels, 2024
Engstrom, L., Feldmann, A., and Madry, A · 2024
Closest in time.