Fetching the paper…
Reading the bibliography…
There is a rapidly growing number of open-source Large Language Models (LLMs) and benchmark datasets to compare them.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. (2019) · 1907
Earlier work this paper cites.
Nearest neighbor pattern classification
Cover, T. and Hart, P. (1967) · 1967
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C. M. and Nasrabadi, N. M. (2006) · 2006
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Hastie, T., Tibshirani, R., Friedman, J. H., and Friedman, J. H. (2009) · 2009
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2020) · 2009
Earlier work this paper cites.
Active learning literature survey
Settles, B. (2009) · 2009
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Bojar, O., Buck, C., Federmann, C., Haddow, B., Koehn, P., Leveling, J., Monz, C., Pecina, P., Post, M., Saint-Amand, H., et al. (2014) · 2014
Earlier work this paper cites.
Deep coral: Correlation alignment for deep domain adaptation
Sun, B. and Saenko, K. (2016) · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017) · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Earlier work this paper cites.
Model evaluation, model selection, and algorithm selection in machine learning
Raschka, S. (2018) · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. (2018) · 2018
Earlier work this paper cites.
Scaling and benchmarking self-supervised visual representation learning
Goyal, P., Mahajan, D., Gupta, A., and Misra, I. (2019) · 2019
Earlier work this paper cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J., Lakshminarayanan, B., and Snoek, J. (2019) · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. (2019) · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I. (2019) · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. (2019) · 2019
Cited alongside, same era.
Frugalml: How to use ml prediction apis more accurately and cheaply
Chen, L., Zaharia, M., and Zou, J. Y. (2020) · 2020
Cited alongside, same era.
Masked language model scoring
Salazar, J., Liang, D., Nguyen, T. Q., and Kirchhoff, K. (2020) · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Sellam, T., Das, D., and Parikh, A. (2020) · 2020
Leveraging unlabeled data to predict out-of-distribution performance
Garg, S., Balakrishnan, S., Lipton, Z. C., Neyshabur, B., and Sedghi, H. (2022) · 2022
Later among the works it cites.
Elevater: A benchmark and toolkit for evaluating language-augmented visual models
Li, C., Liu, H., Li, L., Zhang, P., Aneja, J., Yang, J., Jin, P., Hu, H., Liu, Z., Lee, Y. J., et al. (2022) · 2022
Later among the works it cites.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al. (2022) · 2022
Later among the works it cites.
Summareranker: A multi-task mixture-of-experts re-ranking framework for abstractive summarization
Ravaut, M., Joty, S., and Chen, N. (2022) · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. (2022) · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BERTScore: Evaluating Text Generation with BERT
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. (2020) · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. (2021) · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. (2021) · 2021
Cited alongside, same era.
Predicting with confidence on unseen distributions
Guillory, D., Shankar, V., Ebrahimi, S., Darrell, T., and Schmidt, L. (2021) · 2021
Cited alongside, same era.
In search of lost domain generalization
Gulrajani, I. and Lopez-Paz, D. (2021) · 2021
Cited alongside, same era.
Assessing Generalization of SGD via Disagreement
Jiang, Y., Nagarajan, V., Baek, C., and Kolter, J. Z. (2021) · 2021
Cited alongside, same era.
Later among the works it cites.
Estimation of prediction error with known covariate shift
Xu, H. and Tibshirani, R. (2022) · 2022
Later among the works it cites.
Predicting out-of-distribution error with the projection norm
Yu, Y., Yang, Z., Wei, A., Ma, Y., and Steinhardt, J. (2022) · 2022
Later among the works it cites.
Open llm leaderboard
Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T. (2023) · 2023
Closest in time.
FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
Chen, L., Zaharia, M., and Zou, J. (2023) · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. (2023) · 2023
Closest in time.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
Jiang, D., Ren, X., and Lin, B. Y. (2023) · 2023
Closest in time.
The UCI Machine Learning Repository
Kelly, M., Longjohn, R., and Nottingham, K. (2023) · 2023
Closest in time.
Open assistant
LAION-AI (2023) · 2023
Closest in time.
Understanding new tasks through the lens of training data via exponential tilting
Maity, S., Yurochkin, M., Banerjee, M., and Sun, Y. (2023) · 2023
Closest in time.
Predicting out-of-domain generalization with neighborhood invariance
Ng, N. H., Hulkund, N., Cho, K., and Ghassemi, M. (2023) · 2023
Closest in time.