Fetching the paper…
Reading the bibliography…
Recent work has demonstrated substantial gains in pre-training large-language models (LLMs) followed by supervised fine-tuning on the downstream task.
Brown, Tom, et al. "Language models are few-shot learners." Advances in neural information processing systems 33 (2020): 1877-1901
1901
Earlier work this paper cites.
2016
Earlier work this paper cites.
Vaswani, Ashish, et al. "Attention is all you need." Advances in neural information processing systems 30 (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Poggio, Tomaso, et al. "Why and when can deep-but not shallow-networks avoid the curse of dimensionality: a review." International Journal of Automation and Computing 14.5 (2017): 503-519
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Chollet, François. "On the measure of intelligence." arXiv preprint arXiv:1911.01547 (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Radford, Alec, et al. "Language models are unsupervised multitask learners." OpenAI blog 1.8 (2019): 9
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
Bisk, Yonatan, et al. "Piqa: Reasoning about physical commonsense in natural language." Proceedings of the AAAI conference on artificial intelligence. Vol. 34. No. 05. 2020
2020
Earlier work this paper cites.
Zhou, Xuhui, et al. "Evaluating commonsense in pre-trained language models." Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 34. No. 05. 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Sagawa, Shiori, et al. "An investigation of why overparameterization exacerbates spurious correlations." International Conference on Machine Learning. PMLR, 2020
2020
Cited alongside, same era.
Raffel, Colin, et al. "Exploring the limits of transfer learning with a unified text-to-text transformer." The Journal of Machine Learning Research 21.1 (2020): 5485-5551
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Zhou, Xuhui, et al. "Evaluating commonsense in pre-trained language models." Proceedings of the AAAI conference on artificial intelligence. Vol. 34. No. 05. 2020
2020
Cited alongside, same era.
Sakaguchi, Keisuke, et al. "Winogrande: An adversarial winograd schema challenge at scale." Communications of the ACM 64.9 (2021): 99-106
Li, Xiang Lorraine, et al. "A systematic investigation of commonsense knowledge in large language models." Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Nakkiran, Preetum, et al. "Deep double descent: Where bigger models and more data hurt." Journal of Statistical Mechanics: Theory and Experiment 2021.12 (2021): 124003
2021
Cited alongside, same era.
Bubeck, Sébastien, and Mark Sellke. "A universal law of robustness via isoperimetry." Advances in Neural Information Processing Systems 34 (2021): 28811-28822
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Geva, Mor, et al. "Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies." Transactions of the Association for Computational Linguistics 9 (2021): 346-361
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Elhage, Nelson, et al. "A mathematical framework for transformer circuits." Transformer Circuits Thread 1 (2021)
2021
Cited alongside, same era.
2022
Closest in time.
2022
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Von Oswald, Johannes, et al. "Transformers learn in-context by gradient descent." International Conference on Machine Learning. PMLR, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Refinetti, Maria, Alessandro Ingrosso, and Sebastian Goldt. "Neural networks trained with SGD learn distributions of increasing complexity." International Conference on Machine Learning. PMLR, 2023
2023
Closest in time.
Mitchell, Melanie, and David C. Krakauer. "The debate over understanding in AI’s large language models." Proceedings of the National Academy of Sciences 120.13 (2023)
2023
Closest in time.
Mitchell, Melanie. "How do we know how smart AI systems are?." Science 381.6654 (2023)
2023
Closest in time.