Fetching the paper…
Reading the bibliography…
We present HardML, a benchmark designed to evaluate the knowledge and reasoning abilities in the fields of data science and machine learning.
Brown, T. B., Mann, B., Ryder, N., et al. (2020). Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems, 33, 1877–1901
1901
Earlier work this paper cites.
1904
Earlier work this paper cites.
1904
Earlier work this paper cites.
Bishop, C. M. (2006). Pattern Recognition and Machine Learning. Springer
2006
Earlier work this paper cites.
Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer
2009
Earlier work this paper cites.
Provost, F., & Fawcett, T. (2013). Data Science and its Relationship to Big Data and Data-Driven Decision Making. Big Data, 1(1), 51–59
2013
Earlier work this paper cites.
Jordan, M. I., & Mitchell, T. M. (2015). Machine Learning: Trends, Perspectives, and Prospects. Science, 349(6245), 255–260
2015
Earlier work this paper cites.
Wang, A., Singh, A., Michael, J., et al. (2018). GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In Proceedings of the EMNLP Workshop. Association for Computational Linguistics
2018
Earlier work this paper cites.
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT. Association for Computational Linguistics
2019
Earlier work this paper cites.
Saxton, D., Grefenstette, E., Hill, F., & Kohli, P. (2019). Analysing Mathematical Reasoning Abilities of Neural Models. In International Conference on Learning Representations (ICLR)
2019
Cited alongside, same era.
Raffel, C., Shazeer, N., Roberts, A., et al. (2020). Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research, 21(140), 1–67
2020
Cited alongside, same era.
Hendrycks, D., Burns, C., Basart, S., et al. (2021). Measuring Massive Multitask Language Understanding. In Proceedings of the International Conference on Learning Representations (ICLR)
2021
Cited alongside, same era.
Hendrycks, D., Burns, C., Kadavath, S., et al. (2021). Measuring Mathematical Problem Solving with the MATH Dataset. In Advances in Neural Information Processing Systems
2021
Cited alongside, same era.
OpenAI. (2024). GPT-4o System Card. Retrieved from https://cdn.openai.com/gpt-4o-system-card.pdf
2024
Later among the works it cites.
Anthropic. (2024). Introducing Claude. Retrieved from https://www.anthropic.com/news/introducing-claude
2024
Later among the works it cites.
OpenAI. (2024). Hello GPT-4o. Retrieved from https://openai.com/index/hello-gpt-4o/
2024
Later among the works it cites.
OpenAI. (2024). Introducing OpenAI o1. Retrieved from https://openai.com/index/introducing-openai-o1-preview/?utm_source=chatgpt.com
2024
Later among the works it cites.
OpenAI. (2024). GPT-4o Mini: Advancing Cost-Efficient Intelligence. Retrieved from https://openai.com/blog/gpt-4o-mini-advancing-cost-efficient-intelligence
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dodge, J., Ilharco, G., Schwartz, R., et al. (2021). Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. In Proceedings of the 2021 EMNLP Workshop on Datasets and Benchmarks. Association for Computational Linguistics
2021
Cited alongside, same era.
2022
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
(2024c). Learning to Reason with LLMs. URL: https://openai.com/index/learning-to-reason-withllms/
Cited in the paper.
Meta AI. (2024). Introducing Meta Llama 3: The most capable openly available LLM to date. Retrieved from https://ai.meta.com/blog/llama-3/
2024
Later among the works it cites.
2024
Later among the works it cites.