Fetching the paper…
Reading the bibliography…
Assessing the capabilities and limitations of large language models (LLMs) has garnered significant interest, yet the evaluation of multiple models in real-world scenarios remains rare.
Cross-lingual Name Tagging and Linking for 282 Languages
Pan, X.; Zhang, B.; May, J.; Nothman, J.; Knight, K.; and Ji, H. 2017 · 1958
Earlier work this paper cites.
Health and Health Seeking Behaviour Among Tribal Communities in India: A Socio-Cultural Perspective
Islary, J. 2014 · 2014
Earlier work this paper cites.
Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science
Bender, E. M.; and Friedman, B. 2018 · 2018
Earlier work this paper cites.
XNLI: Evaluating Cross-lingual Sentence Representations
Conneau, A.; Rinott, R.; Lample, G.; Williams, A.; Bowman, S.; Schwenk, H.; and Stoyanov, V. 2018 · 2018
Earlier work this paper cites.
The gendered experience with respect to health-seeking behaviour in an urban slum of Kolkata, India
Das, M.; Angeli, F.; Krumeich, A. J. S. M.; and van Schayck, C. P. 2018 · 2018
Earlier work this paper cites.
Feedpal: Understanding opportunities for chatbots in breastfeeding education of women in India
Yadav, D.; Malik, P.; Dabas, K.; Singh, P.; Deepika YadavIndraprastha Institute of Information Technology, D.; Prerna MalikIndraprastha Institute of Information Technology, D.; Kirti DabasIndraprastha Institute of Information Technology, D.; and Pushpendra SinghIndraprastha Institute of Information Technology, D. 2019 · 2019
Earlier work this paper cites.
Dense Passage Retrieval for Open-Domain Question Answering
Karpukhin, V.; Oguz, B.; Min, S.; Lewis, P.; Wu, L.; Edunov, S.; Chen, D.; and Yih, W.-t. 2020 · 2020
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-t.; Rocktäschel, T.; Riedel, S.; and Kiela, D. 2020 · 2020
Earlier work this paper cites.
No Language Left Behind: Scaling Human-Centered Machine Translation
Team, N.; Costa-jussà, M. R.; Cross, J.; Çelebi, O.; Elbayad, M.; Heafield, K.; Heffernan, K.; Kalbassi, E.; Lam, J.; Licht, D.; Maillard, J.; Sun, A.; Wang, S.; Wenzek, G.; Youngblood, A.; Akula, B.; Barrault, L.; Gonzalez, G. M.; Hansanti, P.; Hoffman, J.; Jarrett, S.; Sadagopan, K. R.; Rowe, D.; Spruit, S.; Tran, C.; Andrews, P.; Ayan, N. F.; Bhosale, S.; Edunov, S.; Fan, A.; Gao, C.; Goswami, V.; Guzmán, F.; Koehn, P.; Mourachko, A.; Ropers, C.; Saleem, S.; Schwenk, H.; and Wang, J. 2022 · 2022
Earlier work this paper cites.
An Artificial Intelligence Chatbot for Young People’s Sexual and Reproductive Health in India (SnehAI): Instrumental Case Study
Wang, H.; Gupta, S.; Singhal, A.; Muttreja, P.; Singh, S.; Sharma, P.; and Piterova, A. 2022 · 2022
Earlier work this paper cites.
MEGA: Multilingual Evaluation of Generative AI
Ahuja, K.; Diddee, H.; Hada, R.; Ochieng, M.; Ramesh, K.; Jain, P.; Nambi, A.; Ganu, T.; Segal, S.; Ahmed, M.; Bali, K.; and Sitaram, S. 2023 · 2023
Earlier work this paper cites.
IndicTrans2: Towards High-Quality and Accessible Machine Translation Models for all 22 Scheduled Indian Languages
Gala, J.; Chitale, P. A.; Raghavan, A. K.; Gumma, V.; Doddapaneni, S.; M, A. K.; Nawale, J. A.; Sujatha, A.; Puduppully, R.; Raghavan, V.; Kumar, P.; Khapra, M. M.; Dabre, R.; and Kunchukuttan, A. 2023 · 2023
Earlier work this paper cites.
Large Language Models Are State-of-the-Art Evaluators of Translation Quality
Kocmi, T.; and Federmann, C. 2023 · 2023
Earlier work this paper cites.
ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning
Lai, V. D.; Ngo, N.; Pouran Ben Veyseh, A.; Man, H.; Dernoncourt, F.; Bui, T.; and Nguyen, T. H. 2023 · 2023
Earlier work this paper cites.
BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models
Leong, W. Q.; Ngui, J. G.; Susanto, Y.; Rengarajan, H.; Sarveswaran, K.; and Tjhi, W. C. 2023 · 2023
Earlier work this paper cites.
Holistic Evaluation of Language Models
Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; Newman, B.; Yuan, B.; Yan, B.; Zhang, C.; Cosgrove, C. A.; Manning, C. D.; Re, C.; Acosta-Navas, D.; Hudson, D. A.; Zelikman, E.; Durmus, E.; Ladhak, F.; Rong, F.; Ren, H.; Yao, H.; WANG, J.; Santhanam, K.; Orr, L.; Zheng, L.; Yuksekgonul, M.; Suzgun, M.; Kim, N.; Guha, N.; Chatterji, N. S.; Khattab, O.; Henderson, P.; Huang, Q.; Chi, R. A.; Xie, S. M.; Santurkar, S.; Ganguli, S.; Hashimoto, T.; Icard, T.; Zhang, T.; Chaudhary, V.; Wang, W.; Li, X.; Mai, Y.; Zhang, Y.; and Koreeda, Y. 2023 · 2023
Earlier work this paper cites.
Hindi Chatbot for Supporting Maternal and Child Health Related Queries in Rural India
Mishra, R.; Singh, S.; Kaur, J.; Singh, P.; and Shah, R. 2023 · 2023
Earlier work this paper cites.
ChatGPT MT: Competitive for High- (but Not Low-) Resource Languages
Robinson, N.; Ogayo, P.; Mortensen, D. R.; and Neubig, G. 2023 · 2023
Cited alongside, same era.
Large Language Models are Not Yet Human-Level Evaluators for Abstractive Summarization
Shen, C.; Cheng, L.; Nguyen, X.-P.; You, Y.; and Bing, L. 2023 · 2023
Cited alongside, same era.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Srivastava, A.; and Team, B.-B. 2023 · 2023
Cited alongside, same era.
Powering an AI Chatbot with Expert Sourcing to Support Credible Health Information Access
Xiao, Z.; Liao, Q. V.; Zhou, M.; Grandison, T.; and Li, Y. 2023 · 2023
Cited alongside, same era.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; Zhang, H.; Gonzalez, J. E.; and Stoica, I. 2023 · 2023
Cited alongside, same era.
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Kim, S.; Suk, J.; Longpre, S.; Lin, B. Y.; Shin, J.; Welleck, S.; Neubig, G.; Lee, M.; Lee, K.; and Seo, M. 2024b · 2024
Closest in time.
OMGEval: An Open Multilingual Generative Evaluation Benchmark for Large Language Models
Liu, Y.; Xu, M.; Wang, S.; Yang, L.; Wang, H.; Liu, Z.; Kong, C.; Chen, Y.; Liu, Y.; Sun, M.; and Yang, E. 2024 · 2024
Closest in time.
Proving Test Set Contamination in Black-Box Language Models
Oren, Y.; Meister, N.; Chatterji, N. S.; Ladhak, F.; and Hashimoto, T. 2024 · 2024
Closest in time.
Evaluating Retrieval Quality in Retrieval-Augmented Generation
Salemi, A.; and Zamani, H. 2024 · 2024
Closest in time.
MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries
Tang, Y.; and Yang, Y. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
Ahuja, S.; Aggarwal, D.; Gumma, V.; Watts, I.; Sathe, A.; Ochieng, M.; Hada, R.; Jain, P.; Ahmed, M.; Bali, K.; and Sitaram, S. 2024 · 2024
Cited alongside, same era.
BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer
Asai, A.; Kudugunta, S.; Yu, X.; Blevins, T.; Gonen, H.; Reid, M.; Tsvetkov, Y.; Ruder, S.; and Hajishirzi, H. 2024 · 2024
Cited alongside, same era.
Benchmarking Large Language Models in Retrieval-Augmented Generation
Chen, J.; Lin, H.; Han, X.; and Sun, L. 2024 · 2024
Cited alongside, same era.
Retrieval-augmented generation in multilingual settings
Chirkova, N.; Rau, D.; Déjean, H.; Formal, T.; Clinchant, S.; and Nikoulina, V. 2024 · 2024
Cited alongside, same era.
Investigating Data Contamination in Modern Benchmarks for Large Language Models
Deng, C.; Zhao, Y.; Tang, X.; Gerstein, M.; and Cohan, A. 2024 · 2024
Cited alongside, same era.
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
Doddapaneni, S.; Khan, M. S. U. R.; Verma, S.; and Khapra, M. M. 2024 · 2024
Cited alongside, same era.
RAGAs: Automated Evaluation of Retrieval Augmented Generation
Es, S.; James, J.; Espinosa Anke, L.; and Schockaert, S. 2024 · 2024
Cited alongside, same era.
PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data
Watts, I.; Gumma, V.; Yadavalli, A.; Seshadri, V.; Swaminathan, M.; and Sitaram, S. 2024 · 2024
Closest in time.
Benchmarking Retrieval-Augmented Generation for Medicine
Xiong, G.; Jin, Q.; Lu, Z.; and Zhang, A. 2024a · 2024
Closest in time.
Benchmark Data Contamination of Large Language Models: A Survey
Xu, C.; Guan, S.; Greene, D.; and Kechadi, M.-T. 2024 · 2024
Closest in time.
CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs’ Cultural Knowledge Through Human-AI Red-Teaming
Chiu, Y. Y.; Jiang, L.; Lin, B. Y.; Park, C. Y.; Li, S. S.; Ravi, S.; Bhatia, M.; Antoniak, M.; Tsvetkov, Y.; Shwartz, V.; and Choi, Y. 2025 · 2025
Closest in time.
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
Doddapaneni, S.; Khan, M. S. U. R.; Venkatesh, D.; Dabre, R.; Kunchukuttan, A.; and Khapra, M. M. 2025 · 2025
Closest in time.
Towards Inducing Long-Context Abilities in Multilingual Neural Machine Translation Models
Gumma, V.; Chitale, P. A.; and Bali, K. 2025 · 2025
Closest in time.
PiCO: Peer Review in LLMs based on Consistency Optimization
Ning, K.-P.; Yang, S.; Liu, Y.; Yao, J.-Y.; Liu, Z.; Tian, Y.; Song, Y.; and Yuan, L. 2025 · 2025
Closest in time.
Beyond Metrics: Evaluating LLMs Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios
Ochieng, M.; Gumma, V.; Sitaram, S.; Wang, J.; Chaudhary, V.; Ronen, K.; Bali, K.; and O’Neill, J. 2025 · 2025
Closest in time.
M-Prometheus: A Suite of Open Multilingual LLM Judges
Pombal, J.; Yoon, D.; Fernandes, P.; Wu, I.; Kim, S.; Rei, R.; Neubig, G.; and Martins, A. 2025 · 2025
Closest in time.
CataractBot: An LLM-powered Expert-in-the-Loop Chatbot for Cataract Patients
Ramjee, P.; Sachdeva, B.; Golechha, S.; Kulkarni, S.; Fulari, G.; Murali, K.; and Jain, M. 2025 · 2025
Closest in time.
Evaluation of Retrieval-Augmented Generation: A Survey
Yu, H.; Gan, A.; Zhang, K.; Tong, S.; Liu, Q.; and Liu, Z. 2025 · 2025
Closest in time.