Fetching the paper…
Reading the bibliography…
We present a comprehensive evaluation of proprietary and open-weights large language models using the first astronomy-specific benchmarking dataset.
1903
Earlier work this paper cites.
1909
Earlier work this paper cites.
http://www.jstor.org/stable/2276774
Wilson, E. B. 1927, Journal of the American Statistical Association, 22, 209 · 1927
Earlier work this paper cites.
2001
Earlier work this paper cites.
Hu, W., & Dodelson, S. 2002, ARA&A, 40, 171, doi: 10.1146/annurev.astro.40.060401.093926
2002
Earlier work this paper cites.
Hendrycks, D., Burns, C., Basart, S., et al. 2020, arXiv preprint arXiv:2009.03300
2009
Earlier work this paper cites.
2011
Earlier work this paper cites.
Bowman, S. R., Angeli, G., Potts, C., & Manning, C. D. 2015, in Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, ed. L. Màrquez, C. Callison-Burch, & J. Su (Lisbon, Portugal: Association for Computational Linguistics), 632–642, doi: 10.18653/v1/D15-1075
2015
Earlier work this paper cites.
Abbott, B. P., Abbott, R., Abbott, T. D., et al. 2016, Phys. Rev. Lett., 116, 061102, doi: 10.1103/PhysRevLett.116.061102
2016
Earlier work this paper cites.
Aihara, H., Armstrong, R., Bickerton, S., et al. 2018, PASJ, 70, S8, doi: 10.1093/pasj/psx081
2018
Earlier work this paper cites.
Chen, D. 2018, PhD thesis, Stanford, CA, USA
2018
Earlier work this paper cites.
Beltagy, I., Lo, K., & Cohan, A. 2019, in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), ed. K. Inui, J. Jiang, V. Ng, & X. Wan (Hong Kong, China: Association for Computational Linguistics), 3615–3620, doi: 10.18653/v1/D19-1371
2019
Earlier work this paper cites.
Wang, A., Pruksachatkun, Y., Nangia, N., et al. 2019, Advances in neural information processing systems, 32
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., et al. 2020, Advances in neural information processing systems, 33, 1877
2020
Earlier work this paper cites.
Lin, S., Hilton, J., & Evans, O. 2021, arXiv preprint arXiv:2109.07958
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
Liang, P., Bommasani, R., Lee, T., et al. 2022, arXiv preprint arXiv:2211.09110
2022
Earlier work this paper cites.
Luo, L., Lai, P.-T., Wei, C.-H., Arighi, C. N., & Lu, Z. 2022, Briefings in Bioinformatics, 23, bbac282, doi: 10.1093/bib/bbac282
2022
Earlier work this paper cites.
Saikh, T., Ghosal, T., Mittal, A., Ekbal, A., & Bhattacharyya, P. 2022, International Journal on Digital Libraries, 23, 289, doi: 10.1007/s00799-022-00329-y
2022
Cited alongside, same era.
Wang, Y., Zhai, Z., Alavi, A., et al. 2022, ApJ, 928, 1, doi: 10.3847/1538-4357/ac4973
2022
Cited alongside, same era.
2022
Cited alongside, same era.
https://research.google/pubs/archive/chain-of-thought/
Zhou, D., & Wei, J. 2022, Google Research Blog · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., et al. 2023, arXiv preprint arXiv:2303.08774
2023
Cited alongside, same era.
https://openreview.net/forum?id=uccHPGDlao
Zheng, L., Chiang, W.-L., Sheng, Y., et al. 2023, in Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track · 2023
Later among the works it cites.
2024
Closest in time.
https://arxiv.org/abs/2403.17297
Cai, Z., Cao, M., Chen, H., et al. 2024, InternLM2 Technical Report · 2024
Closest in time.
2024
Closest in time.
https://arxiv.org/abs/2403.04132
Chiang, W.-L., Zheng, L., Sheng, Y., et al. 2024, Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Blecher, L., Cucurull, G., Scialom, T., & Stojnic, R. 2023, arXiv preprint arXiv:2308.13418
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Gao, L., Tow, J., Abbasi, B., et al. 2023, A framework for few-shot language model evaluation, v0.4.0, Zenodo, doi: 10.5281/zenodo.10256836
2023
Cited alongside, same era.
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Perkowski, E., Pan, R., Nguyen, T. D., et al. 2024, Research Notes of the American Astronomical Society, 8, 7, doi: 10.3847/2515-5172/ad1abe
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
https://openreview.net/forum?id=BOfDKxfwt0
Zheng, L., Chiang, W.-L., Sheng, Y., et al. 2024, in The Twelfth International Conference on Learning Representations · 2024
Closest in time.