Fetching the paper…
Reading the bibliography…
We introduce the Falcon series: 7B, 40B, and 180B parameters causal decoder-only models trained on a diverse high-quality corpora predominantly assembled from web data.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 1901
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. (2019) · 1907
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. (2019) · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2019) · 1910
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Shazeer, N. (2019) · 1911
Earlier work this paper cites.
Prediction and entropy of printed english
Shannon, C. E. (1951) · 1951
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Suffix arrays: a new method for on-line string searches
Manber, U. and Myers, G. (1993) · 1993
Earlier work this paper cites.
On the resemblance and containment of documents
Broder, A. Z. (1997) · 1997
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Adiwardana, D., Luong, M.-T., So, D. R., Hall, J., Fiedel, N., Thoppilan, R., Yang, Z., Kulshreshtha, A., Nemade, G., Lu, Y., et al. (2020) · 2001
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2001
Earlier work this paper cites.
Glu variants improve transformer
Shazeer, N. (2020) · 2002
Earlier work this paper cites.
Gromov–wasserstein distances and the metric approach to object matching
Mémoli, F. (2011) · 2011
Earlier work this paper cites.
SemEval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Gordon, A., Kozareva, Z., and Roemmele, M. (2012) · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T. (2013) · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013) · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D. (2014) · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X. (2015) · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. (2015) · 2015
Earlier work this paper cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Inan, H., Khosravi, K., and Socher, R. (2016) · 2016
Earlier work this paper cites.
Exploring the limits of language modeling
Jozefowicz, R., Vinyals, O., Schuster, M., Shazeer, N., and Wu, Y. (2016) · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N. Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R. (2016) · 2016
Earlier work this paper cites.
Using the output embedding to improve language models
Press, O. and Wolf, L. (2016) · 2016
Earlier work this paper cites.
Finding alternative translations in a large corpus of movie subtitle
Tiedemann, J. (2016) · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al. (2016) · 2016
Earlier work this paper cites.
news-please: A generic news crawler and extractor
Hamborg, F., Meuschke, N., Breitinger, C., and Gipp, B. (2017) · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y. (2017) · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Johannes Welbl, Nelson F. Liu, M. G. (2017) · 2017
Earlier work this paper cites.
Race: Large-scale reading comprehension dataset from examinations
Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E. (2017) · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2017) · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the AI2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. (2018) · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S. (2018) · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2018) · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. (2018) · 2018
Earlier work this paper cites.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018) · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2018) · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., and Vaswani, A. (2018) · 2018
Earlier work this paper cites.
Mesh-TensorFlow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., Sepassi, R., and Hechtman, B. (2018) · 2018
Earlier work this paper cites.
A simple method for commonsense reasoning
Trinh, T. H. and Le, Q. V. (2018) · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. (2018) · 2018
Earlier work this paper cites.
Scibert: A pretrained language model for scientific text
Beltagy, I., Lo, K., and Cohan, A. (2019) · 2019
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K. (2019) · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Kenton, J. D. M.-W. C. and Toutanova, L. K. (2019) · 2019
Earlier work this paper cites.
Model cards for model reporting
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., and Gebru, T. (2019) · 2019
Earlier work this paper cites.
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures
Ortiz Suárez, P. J., Sagot, B., and Romary, L. (2019) · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. (2019) · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., and Lillicrap, T. P. (2019) · 2019
Earlier work this paper cites.
The bitter lesson
Sutton, R. (2019) · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Tillet, P., Kung, H.-T., and Cox, D. (2019) · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. (2019) · 2019
Cited alongside, same era.
HellaSwag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 2019
Cited alongside, same era.
The pushshift reddit dataset
Baumgartner, J., Zannettou, S., Keegan, B., Squire, M., and Blackburn, J. (2020) · 2020
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y. (2020) · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Deduplicating training data makes language models better
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N. (2022) · 2022
Later among the works it cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y., Neyshabur, B., Gur-Ari, G., and Misra, V. (2022) · 2022
Later among the works it cites.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al. (2022) · 2022
Later among the works it cites.
Your transformer may not be as powerful as you expect
Luo, S., Li, S., Zheng, S., Liu, T.-Y., Wang, L., and He, D. (2022) · 2022
Later among the works it cites.
Language models of code are few-shot commonsense learners
Madaan, A., Zhou, S., Alon, U., Yang, Y., and Neubig, G. (2022) · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, É., Ott, M., Zettlemoyer, L., and Stoyanov, V. (2020) · 2020
Cited alongside, same era.
The Pile: an 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C. (2020) · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2020) · 2020
Cited alongside, same era.
Limits to depth efficiencies of self-attention
Levine, Y., Wies, N., Sharir, O., Bata, H., and Shashua, A. (2020) · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. (2020) · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. (2020) · 2020
Cited alongside, same era.
Ccnet: Extracting high quality monolingual datasets from web crawl data
Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzmán, F., Joulin, A., and Grave, É. (2020) · 2020
Cited alongside, same era.
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N., and Lewis, M. (2022) · 2022
Later among the works it cites.
Scaling up models and data with t5x
Roberts, A., Chung, H. W., Levskaya, A., Mishra, G., Bradbury, J., Andor, D., Narang, S., Lester, B., Gaffney, C., Mohiuddin, A., Hawthorne, C., Lewkowycz, A., Salcianu, A., van Zee, M., Austin, J., Goodman, S., Soares, L. B., Hu, H., Tsvyashchenko, S., Chowdhery, A., Bastings, J., Bulian, J., Garcia, X., Ni, J., Chen, A., Kenealy, K., Clark, J. H., Lee, S., Garrette, D., Lee-Thorp, J., Raffel, C., Shazeer, N., Ritter, M., Bosma, M., Passos, A., Maitin-Shepard, J., Fiedel, N., Omernick, M., Saeta, B., Sepassi, R., Spiridonov, A., Newlan, J., and Gesmundo, A. (2022) · 2022
Later among the works it cites.
Causes and cures for interference in multilingual translation
Shaham, U., Elbayad, M., Goswami, V., Levy, O., and Bhosale, S. (2022) · 2022
Later among the works it cites.
Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., et al. (2022) · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al. (2022) · 2022
Later among the works it cites.
Amos: An adam-style optimizer with adaptive weight decay towards model-oriented scale
Tian, R. and Parikh, A. P. (2022) · 2022
Later among the works it cites.
Will we run out of data? an analysis of the limits of scaling datasets in machine learning
Villalobos, P., Sevilla, J., Heim, L., Besiroglu, T., Hobbhahn, M., and Ho, A. (2022) · 2022
Later among the works it cites.
Byt5: Towards a token-free future with pre-trained byte-to-byte models
Xue, L., Barua, A., Constant, N., Al-Rfou, R., Narang, S., Kale, M., Roberts, A., and Raffel, C. (2022) · 2022
Later among the works it cites.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J. (2022) · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. (2022) · 2022
Later among the works it cites.
Scaling laws for generative mixed-modal language models
Aghajanyan, A., Yu, L., Conneau, A., Hsu, W.-N., Hambardzumyan, K., Zhang, S., Roller, S., Goyal, N., Levy, O., and Zettlemoyer, L. (2023) · 2023
Closest in time.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S. (2023) · 2023
Closest in time.
Luminous: performance benchmarks
Aleph Alpha (2023) · 2023
Closest in time.
Santacoder: don’t reach for the stars!
Allal, L. B., Li, R., Kocetkov, D., Mou, C., Akiki, C., Ferrandis, C. M., Muennighoff, N., Mishra, M., Gu, A., Dey, M., et al. (2023) · 2023
Closest in time.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. (2023) · 2023
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al. (2023) · 2023
Closest in time.
Extending context window of large language models via positional interpolation
Chen, S., Wong, S., Chen, L., and Tian, Y. (2023) · 2023
Closest in time.
Can large language models be an alternative to human evaluations?
Chiang, C.-H. and Lee, H.-y. (2023) · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. (2023) · 2023
Closest in time.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T. (2023) · 2023
Closest in time.
Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster
Dey, N., Gosal, G., Khachane, H., Marshall, W., Pathria, R., Tom, M., Hestness, J., et al. (2023) · 2023
Closest in time.
Effective theory of transformers at initialization
Dinan, E., Yaida, S., and Zhang, S. (2023) · 2023
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Closest in time.
Ethnologue: Languages of the World
Eberhard, D. M., Simons, G. F., and Fennig, C. D. (2023) · 2023
Closest in time.
What’s going on with the open llm leaderboard?
Fourrier, C., Habib, N., Launay, J., and Wolf, T. (2023) · 2023
Closest in time.
Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Del Giorno, A., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., et al. (2023) · 2023
Closest in time.
Inflection 1
Inflection (2023) · 2023
Closest in time.
Extending context is hard… but not impossible
kaiokenmdenv (2023) · 2023
Closest in time.
Xlm-v: Overcoming the vocabulary bottleneck in multilingual masked language models
Liang, D., Gonen, H., Mao, Y., Hou, R., Goyal, N., Ghazvininejad, M., Zettlemoyer, L., and Khabsa, M. (2023) · 2023
Closest in time.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al. (2023) · 2023
Closest in time.
Wizardcoder: Empowering code large language models with evol-instruct
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D. (2023) · 2023
Closest in time.
Introducing mpt-30b: Raising the bar for open-source foundation models
MosaicML (2023) · 2023
Closest in time.
Scaling data-constrained language models
Muennighoff, N., Rush, A. M., Barak, B., Scao, T. L., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T., and Raffel, C. (2023) · 2023
Closest in time.
Model index for researchers
OpenAI (2023b) · 2023
Closest in time.
Stack overflow will charge ai giants for training data
Paresh, D. (2023) · 2023
Closest in time.
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J. (2023) · 2023
Closest in time.
Efficiently scaling transformer inference
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Heek, J., Xiao, K., Agrawal, S., and Dean, J. (2023) · 2023
Closest in time.
Code llama: Open foundation models for code
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al. (2023) · 2023
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. (2023) · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Closest in time.
Large language models are not fair evaluators
Wang, P., Li, L., Chen, L., Zhu, D., Lin, B., Cao, Y., Liu, Q., Liu, T., and Sui, Z. (2023) · 2023
Closest in time.
Effective long-context scaling of foundation models
Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H. (2023) · 2023
Closest in time.
To repeat or not to repeat: Insights from scaling llm under token-crisis
Xue, F., Fu, Y., Zhou, W., Zheng, Z., and You, Y. (2023) · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. (2023) · 2023
Closest in time.
Agieval: A human-centric benchmark for evaluating foundation models
Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N. (2023) · 2023
Closest in time.