Fetching the paper…
Reading the bibliography…
This report evaluates the performance of text-in text-out Large Language Models (LLMs) to understand and generate Indic languages.
PMIndia – A Collection of Parallel Corpora of Languages of India
Barry Haddow and Faheem Kirefu. 2020 · 2001
Earlier work this paper cites.
X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language Models
Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, and Graham Neubig. 2020 · 2010
Earlier work this paper cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2010
Earlier work this paper cites.
Paper 1 of 2018 on Language, Table C-16, Census of India 2011
2011 · 2011
Earlier work this paper cites.
Indic-Transformers: An Analysis of Transformer Language Models for Indian Languages
Kushal Jain, Adwait Deshpande, Kumar Shridhar, Felix Laumann, and Ayushman Dash. 2020 · 2011
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
A Conneau. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
The Language Geography of Wikipedia
Martin Dittus and Mark Graham. 2019 · 2019
Earlier work this paper cites.
IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages. In Findings of the Association for Computational Linguistics: EMNLP 2020 . 4948–4961
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, NC Gokul, Avik Bhattacharyya, Mitesh M Khapra, and Pratyush Kumar. 2020 · 2020
Earlier work this paper cites.
GLUECoS: An evaluation benchmark for code-switched NLP
Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram, and Monojit Choudhury. 2020 · 2020
Earlier work this paper cites.
OdiEnCorp 2.0: Odia-English Parallel Corpus for Machine Translation. In Proceedings of the WILDRE5– 5th Workshop on Indian Language Data: Resources and Evaluation , Girish Nath Jha, Kalika Bali, Sobha L., S. S. Agrawal, and Atul Kr. Ojha (Eds.). European Language Resources Association (ELRA), Marseille, France, 14–19
Shantipriya Parida, Satya Ranjan Dash, Ondřej Bojar, Petr Motlicek, Priyanka Pattnaik, and Debasish Kumar Mallick. 2020 · 2020
Earlier work this paper cites.
Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset. In Proceedings of the Twelfth Language Resources and Evaluation Conference , Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk, and Stelios Piperidis (Eds.). European Language Resources Association, Marseille, France, 2413–2423
Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke, Cibu Johny, Isin Demirsahin, and Keith Hall. 2020 · 2020
Earlier work this paper cites.
Are All Languages Created Equal in Multilingual BERT?. In Proceedings of the 5th Workshop on Representation Learning for NLP , Spandana Gella, Johannes Welbl, Marek Rei, Fabio Petroni, Patrick Lewis, Emma Strubell, Minjoon Seo, and Hannaneh Hajishirzi (Eds.). Association for Computational Linguistics, Online, 120–130
Shijie Wu and Mark Dredze. 2020 · 2020
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
L Xue. 2020 · 2020
Earlier work this paper cites.
How Linguistically Fair are Multilingual Pre-Trained Language Models?. In AAAI-21 . AAAI, AAAI
Monojit Choudhury and Amit Deshpande. 2021 · 2021
Earlier work this paper cites.
IndicBART: A pre-trained model for indic natural language generation
Raj Dabre, Himani Shrotriya, Anoop Kunchukuttan, Ratish Puduppully, Mitesh M Khapra, and Pratyush Kumar. 2021 · 2021
Earlier work this paper cites.
XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages
Tahmid Hasan, Abhik Bhattacharjee, Md Saiful Islam, Kazi Samin, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021 · 2021
Earlier work this paper cites.
Multilingual LAMA: Investigating Knowledge in Multilingual Pretrained Language Models
Nora Kassner, Philipp Dufter, and Hinrich Schütze. 2021 · 2021
Earlier work this paper cites.
Muril: Multilingual representations for indian languages
Simran Khanuja, Diksha Bansal, Sarvesh Mehtani, Savya Khosla, Atreyee Dey, Balaji Gopalan, Dilip Kumar Margam, Pooja Aggarwal, Rajiv Teja Nagipogu, Shachi Dave, et al · 2021
Earlier work this paper cites.
XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 10215–10245
Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021 · 2021
Earlier work this paper cites.
Predicting the Performance of Multilingual NLP Models
Anirudh Srinivasan, Sunayana Sitaram, Tanuja Ganu, Sandipan Dandapat, Kalika Bali, and Monojit Choudhury. 2021 · 2021
Earlier work this paper cites.
IndicXNLI: Evaluating Multilingual Inference for Indian Languages
Divyanshu Aggarwal, Vivek Gupta, and Anoop Kunchukuttan. 2022 · 2022
Earlier work this paper cites.
HateCheckHIn: Evaluating Hindi Hate Speech Detection Models
Mithun Das, Punyajoy Saha, Binny Mathew, and Animesh Mukherjee. 2022 · 2022
Earlier work this paper cites.
Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M Khapra, Anoop Kunchukuttan, and Pratyush Kumar. 2022 · 2022
Earlier work this paper cites.
Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, et al · 2022
Earlier work this paper cites.
IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages
Aman Kumar, Himani Shrotriya, Prachi Sahu, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, Amogh Mishra, Mitesh M. Khapra, and Pratyush Kumar. 2022 · 2022
Cited alongside, same era.
Naamapadam: a large-scale named entity annotated data for Indic languages
Arnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M Khapra, Pratyush Kumar, Rudra Murthy V, and Anoop Kunchukuttan. 2022 · 2022
Cited alongside, same era.
Universal Dependency Treebank for Odia Language
Shantipriya Parida, Kalyanamalini Sahoo, Atul Kr. Ojha, Saraswati Sahoo, Satya Ranjan Dash, and Bijayalaxmi Dash. 2022 · 2022
Cited alongside, same era.
Samanantar: The largest publicly available parallel corpora collection for 11 indic languages
Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan Ak, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Divyanshu Kakwani, Navneet Kumar, et al · 2022
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
Mixtral of experts
Mistral Team. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
On evaluating and mitigating gender biases in multilingual settings
Aniket Vashishtha, Kabir Ahuja, and Sunayana Sitaram. 2023 · 2023
Later among the works it cites.
Lmsys-chat-1m: A large-scale real-world llm conversation dataset
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Generative Chatbot Adaptation for Odia Language: A Critical Evaluation. In 2023 1st International Conference on Circuits, Power and Intelligent Systems (CCPIS) . 1–7
Parul Agarwal, Aisha Asif, Shantipriya Parida, Sambit Sekhar, Satya Ranjan Dash, and Subhadarshi Panda. 2023 · 2023
Cited alongside, same era.
Mega: Multilingual evaluation of generative ai
Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Maxamed Axmed, et al · 2023
Cited alongside, same era.
Megaverse: Benchmarking large language models across languages, modalities, models and tasks
Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Prachi Jain, Maxamed Axmed, Kalika Bali, et al · 2023
Cited alongside, same era.
Taking Stock of Concept Inventories in Computing Education: A Systematic Literature Review. In Proceedings of the 2023 ACM Conference on International Computing Education Research V.1 (ICER 2023) . ACM, 397–415
Murtaza Ali, Sourojit Ghosh, Prerna Rao, Raveena Dhegaskar, Sophia Jawort, Alix Medler, Mengqi Shi, and Sayamindu Dasgupta. 2023 · 2023
Cited alongside, same era.
Tamil-llama: A new tamil language model based on llama 2
Abhinand Balachandran. 2023 · 2023
Cited alongside, same era.
BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla. In Findings of the Association for Computational Linguistics: EACL 2023 , Andreas Vlachos and Isabelle Augenstein (Eds.). Association for Computational Linguistics, Dubrovnik, Croatia, 726–735
Abhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, and Rifat Shahriyar. 2023 · 2023
Cited alongside, same era.
Vistaar: Diverse Benchmarks and Training Sets for Indian Language ASR
Kaushal Santosh Bhogale, Sai Sundaresan, Abhigyan Raman, Tahir Javed, Mitesh M. Khapra, and Pratyush Kumar. 2023 · 2023
Cited alongside, same era.
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P Xing, et al · 2023
Later among the works it cites.
Rajbhasha Vibhag, Ministry of Home Affairs
2024 · 2024
Later among the works it cites.
Maple: Multilingual evaluation of parameter efficient finetuning of large language models
Divyanshu Aggarwal, Ashutosh Sathe, and Sunayana Sitaram. 2024 · 2024
Later among the works it cites.
Aya 23: Open Weight Releases to Further Multilingual Progress
Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, Kelly Marchisio, Max Bartolo, Sebastian Ruder, Acyr Locatelli, Julia Kreutzer, Nick Frosst, Aidan Gomez, Phil Blunsom, Marzieh Fadaee, Ahmet Üstün, and Sara Hooker. 2024 · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Airavata: Introducing hindi instruction-tuned llm
Jay Gala, Thanmay Jayakumar, Jaavid Aktar Husain, Mohammed Safi Ur Rahman Khan, Diptesh Kanojia, Ratish Puduppully, Mitesh M Khapra, Raj Dabre, Rudra Murthy, Anoop Kunchukuttan, et al · 2024
Later among the works it cites.
METAL: Towards Multilingual Meta-Evaluation
Rishav Hada, Varun Gumma, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2024 · 2024
Later among the works it cites.
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
Carolin Holtermann, Paul Röttger, Timm Dill, and Anne Lauscher. 2024 · 2024
Later among the works it cites.
Mohsinul Kabir, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Mir Tafseer Nayeem, M Saiful Bari, and Enamul Hoque. 2024 · 2024
Later among the works it cites.
Decoding the Diversity: A Review of the Indic AI Research Landscape
Sankalp KJ, Vinija Jain, Sreyoshi Bhaduri, Tamoghna Roy, and Aman Chadha. 2024 · 2024
Later among the works it cites.
The rise of Generative AI large language models (llms) like chatgpt
David McCandless. 2024 · 2024
Later among the works it cites.
Paramanu: A Family of Novel Efficient Indic Generative Foundation Language Models
Mitodru Niyogi and Arnab Bhattacharya. 2024 · 2024
Later among the works it cites.
Millicent Ochieng, Varun Gumma, Sunayana Sitaram, Jindong Wang, Keshet Ronen, Kalika Bali, and Jacki O’Neill. 2024 · 2024
Later among the works it cites.
Hello GPT-4o
OpenAI. 2024 · 2024
Later among the works it cites.
Sarvam 1
Sarvam. 2024 · 2024
Later among the works it cites.
Harman Singh, Nitish Gupta, Shikhar Bharadwaj, Dinesh Tewari, and Partha Talukdar. 2024 · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Later among the works it cites.
MILU: A Multi-task Indic Language Understanding Benchmark
Sshubam Verma, Mohammed Safi Ur Rahman Khan, Vishwajeet Kumar, Rudra Murthy, and Jaydeep Sen. 2024 · 2024
Later among the works it cites.
Ishaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri, Manohar Swaminathan, and Sunayana Sitaram. 2024 · 2024
Later among the works it cites.
A Survey of Large Language Models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2024 · 2024
Later among the works it cites.