Fetching the paper…
Reading the bibliography…
The training of large language models (LLMs) is expensive.
Remarks on some nonparametric estimates of a density function
Rosenblatt, M · 1956
Earlier work this paper cites.
Monte carlo sampling methods using markov chains and their applications
Hastings, W. K · 1970
Earlier work this paper cites.
The equivalence of weak, strong and complete convergence in l1 for kernel density estimates
Devroye, L · 1983
Earlier work this paper cites.
The central role of the propensity score in observational studies for causal effects
Rosenbaum, P. R. and Rubin, D. B · 1983
Earlier work this paper cites.
Locality-sensitive hashing scheme based on p-stable distributions
Datar, M., Immorlica, N., Indyk, P., and Mirrokni, V. S · 2004
Earlier work this paper cites.
Distance-sensitive bloom filters
Kirsch, A. and Mitzenmacher, M · 2006
Earlier work this paper cites.
Super-samples from kernel herding
Chen, Y., Welling, M., and Smola, A · 2012
Earlier work this paper cites.
Consistency of the kernel density estimator: a survey
Wied, D. and Weißbach, R · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T · 2013
Earlier work this paper cites.
Composable core-sets for diversity and coverage maximization
Indyk, P., Mahabadi, S., Mahdian, M., and Mirrokni, V. S · 2014
Earlier work this paper cites.
Generalized outlier detection with flexible kernel density estimates
Schubert, E., Zimek, A., and Kriegel, H.-P · 2014
Earlier work this paper cites.
Teaching machines to read and comprehend
Hermann, K. M., Kocisky, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., and Blunsom, P · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Coresets and sketches
Phillips, J. M · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Not all samples are created equal: Deep learning with importance sampling
Katharopoulos, A. and Fleuret, F · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Accelerating deep learning by focusing on the biggest losers
Jiang, A. H., Wong, D. L.-K., Zhou, G., Andersen, D. G., Dean, J., Ganger, G. R., Joshi, G., Kaminksy, M., Kozuch, M., Lipton, Z. C., et al · 2019
Earlier work this paper cites.
Discrepancy, coresets, and sketches in machine learning
Karnin, Z. and Liberty, E · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Rehashing kernel evaluation in high dimensions
Siminelakis, P., Rong, K., Bailis, P., Charikar, M., and Levis, P · 2019
Earlier work this paper cites.
Energy and policy considerations for deep learning in nlp
Strubell, E., Ganesh, A., and McCallum, A · 2019
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
Toneva, M., Sordoni, A., Combes, R., Trischler, A., Bengio, Y., and Gordon, G · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2019
Earlier work this paper cites.
Ccnet: Extracting high quality monolingual datasets from web crawl data
Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzmán, F., Joulin, A., and Grave, E · 2019
Earlier work this paper cites.
Coresets via bilevel optimization for continual learning and streaming
Borsos, Z., Mutny, M., and Krause, A · 2020
Earlier work this paper cites.
Sub-linear race sketches for approximate kernel density estimation on streaming data
Coleman, B. and Shrivastava, A · 2020
Cited alongside, same era.
Selection via proxy: Efficient data selection for deep learning
Coleman, C., Yeh, C., Mussmann, S., Mirzasoleiman, B., Bailis, P., Liang, P., Leskovec, J., and Zaharia, M · 2020
Cited alongside, same era.
Turning big data into tiny data: Constant-size coresets for k-means, pca, and projective clustering
Feldman, D., Schmidt, M., and Sohler, C · 2020
Cited alongside, same era.
What neural networks memorize and why: Discovering the long tail via influence estimation
Feldman, V. and Zhang, C · 2020
Cited alongside, same era.
Wiki-40b: Multilingual language model dataset
Guo, M., Dai, Z., Vrandečić, D., and Al-Rfou, R · 2020
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.
Semdedup: Data-efficient learning at web-scale through semantic deduplication
Abbas, A., Tirumala, K., Simig, D., Ganguli, S., and Morcos, A. S · 2023
Later among the works it cites.
Gkd: Generalized knowledge distillation for auto-regressive sequence models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Unifiedqa: Crossing format boundaries with a single qa system
Khashabi, D., Min, S., Khot, T., Sabharwal, A., Tafjord, O., Clark, P., and Hajishirzi, H · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Training data subset search with ensemble active learning
Chitta, K., Álvarez, J. M., Haussmann, E., and Farabet, C · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Cited alongside, same era.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Geva, M., Khashabi, D., Segal, E., Khot, T., Roth, D., and Berant, J · 2021
Cited alongside, same era.
Agarwal, R., Vieillard, N., Stanczyk, P., Ramos, S., Geist, M., and Bachem, O · 2023
Later among the works it cites.
Self-consuming generative models go mad
Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A. I., Babaei, H., LeJeune, D., Siahkoohi, A., and Baraniuk, R. G · 2023
Later among the works it cites.
Palm 2 technical report, 2023
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., and et al., Z. C · 2023
Later among the works it cites.
Large language models suffer from their own output: An analysis of the self-consuming training loop
Briesch, M., Sobania, D., and Rothlauf, F · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini, T., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Simfluence: Modeling the influence of individual training examples by simulating training runs
Guu, K., Webson, A., Pavlick, E., Dixon, L., Tenney, I., and Bolukbasi, T · 2023
Later among the works it cites.
Phi-2: The surprising power of small language models, 2023
Javaheripi, M., Bubeck, S., Abdin, M., Aneja, J., Bubeck, S., Mendes, C. C. T., Chen, W., Del Giorno, A., Eldan, R., Gopi, S., et al · 2023
Later among the works it cites.
Lee, A., Miranda, B., and Koyejo, S · 2023
Later among the works it cites.
One-pass distribution sketch for measuring data heterogeneity in federated learning
Liu, Z., Xu, Z., Coleman, B., and Shrivastava, A · 2023
Later among the works it cites.
Estimating the carbon footprint of bloom, a 176b parameter language model
Luccioni, A. S., Viguier, S., and Ligozat, A.-L · 2023
Later among the works it cites.
When less is more: Investigating data pruning for pretraining llms at scale
Marion, M., Üstün, A., Pozzobon, L., Wang, A., Fadaee, M., and Hooker, S · 2023
Later among the works it cites.
Scaling data-constrained language models
Muennighoff, N., Rush, A. M., Barak, B., Scao, T. L., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T., and Raffel, C · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Data distillation: A survey
Sachdeva, N. and McAuley, J · 2023
Later among the works it cites.
Farzi data: Autoregressive data distillation
Sachdeva, N., He, Z., Kang, W.-C., Ni, J., Cheng, D. Z., and McAuley, J · 2023
Later among the works it cites.
Mixture-of-experts meets instruction tuning: A winning combination for large language models
Shen, S., Hou, L., Zhou, Y., Du, N., Longpre, S., Wei, J., Chung, H. W., Zoph, B., Fedus, W., Chen, X., et al · 2023
Later among the works it cites.
The curse of recursion: Training on generated data makes models forget.(5 2023)
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., and Anderson, R · 2023
Later among the works it cites.
D4: Improving llm pretraining via document de-duplication and diversification
Tirumala, K., Simig, D., Aghajanyan, A., and Morcos, A. S · 2023
Later among the works it cites.
Large transformer model inference optimization
Weng, L · 2023
Later among the works it cites.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2023
Later among the works it cites.
Dsdm: Model-aware dataset selection with datamodels, 2024
Engstrom, L., Feldmann, A., and Madry, A · 2024
Closest in time.
Rephrasing the web: A recipe for compute and data-efficient language modeling, 2024
Maini, P., Seto, S., Bai, H., Grangier, D., Zhang, Y., and Jaitly, N · 2024
Closest in time.