Fetching the paper…
Reading the bibliography…
Current Large Language Models (LLMs) are predominantly designed with English as the primary language, and even the few that are multilingual tend to exhibit strong English-centric biases.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Weisfeiler-lehman graph kernels
Nino Shervashidze, Pascal Schweitzer, Erik Jan Van Leeuwen, Kurt Mehlhorn, and Karsten M Borgwardt. 2011 · 2011
Earlier work this paper cites.
A kernel two-sample test
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander Smola. 2012 · 2012
Earlier work this paper cites.
On the features of translationese
Vered Volansky, Noam Ordan, and Shuly Wintner. 2015 · 2015
Earlier work this paper cites.
Crowd-sourcing NLG data: Pictures elicit better data
Jekaterina Novikova, Oliver Lemon, and Verena Rieser. 2016 · 2016
Earlier work this paper cites.
Translationese: Between human and machine translation
Shuly Wintner. 2016 · 2016
Earlier work this paper cites.
Treat the system like a human student: Automatic naturalness evaluation of generated text without reference texts
Isabel Groves, Ye Tian, and Ioannis Douratsos. 2018 · 2018
Earlier work this paper cites.
How human is machine translationese? comparing human and machine translations of text and speech
Yuri Bizzoni, Tom S Juzek, Cristina España-Bonet, Koel Dutta Chowdhury, Josef van Genabith, and Elke Teich. 2020 · 2020
Earlier work this paper cites.
Universal Dependencies v2: An evergrowing multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020 · 2020
Earlier work this paper cites.
Stanza: A python natural language processing toolkit for many human languages
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020 · 2020
Earlier work this paper cites.
Grakel: A graph kernel library in python
Giannis Siglidis, Giannis Nikolentzos, Stratis Limnios, Christos Giatsidis, Konstantinos Skianis, and Michalis Vazirgiannis. 2020 · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Naturalness evaluation of natural language generation in task-oriented dialogues using BERT
Ye Liu, Wolfgang Maier, Wolfgang Minker, and Stefan Ultes. 2021 · 2021
Earlier work this paper cites.
Language model evaluation beyond perplexity
Clara Meister and Ryan Cotterell. 2021 · 2021
Earlier work this paper cites.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021 · 2021
Cited alongside, same era.
Evaluating the evaluation of diversity in natural language generation
Guy Tevet and Jonathan Berant. 2021 · 2021
Cited alongside, same era.
Machine translationese: Effects of algorithmic bias on linguistic complexity in machine translation
Eva Vanmassenhove, Dimitar Shterionov, and Matthew Gwilliam. 2021 · 2021
Cited alongside, same era.
FairLex: A multilingual benchmark for evaluating fairness in legal text processing
Ilias Chalkidis, Tommaso Pasini, Sheng Zhang, Letizia Tomada, Sebastian Schwemer, and Anders Søgaard. 2022 · 2022
Cited alongside, same era.
A natural diet: Towards improving naturalness of machine translation output
Markus Freitag, David Vilar, David Grangier, Colin Cherry, and George Foster. 2022 · 2022
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Later among the works it cites.
Firefly(流萤): 中文对话式大语言模型
Jianxin Yang. 2023 · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al. 2024 · 2024
Closest in time.
Iterative translation refinement with large language models
Pinzhen Chen, Zhicheng Guo, Barry Haddow, and Kenneth Heafield. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
The “problem” of human label variation: On ground truth in data, modeling and evaluation
Barbara Plank. 2022 · 2022
Cited alongside, same era.
Openlabel-chinese conversations dataset
BAAI. 2023 · 2023
Cited alongside, same era.
Assessing syntactic and lexicogrammatical use in second language mandarin writing samples
Susanne DeVore and Kristopher Kyle. 2023 · 2023
Cited alongside, same era.
FactKB: Generalizable factuality evaluation using language models enhanced with factual knowledge
Shangbin Feng, Vidhisha Balachandran, Yuyang Bai, and Yulia Tsvetkov. 2023 · 2023
Cited alongside, same era.
What comes next? evaluating uncertainty in neural text generators against human production variability
Mario Giulianelli, Joris Baan, Wilker Aziz, Raquel Fernández, and Barbara Plank. 2023 · 2023
Cited alongside, same era.
How close is chatgpt to human experts? comparison corpus, evaluation, and detection
Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023 · 2023
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. 2024 · 2024
Closest in time.
The curious decline of linguistic diversity: Training language models on synthetic text
Yanzhu Guo, Guokan Shang, Michalis Vazirgiannis, and Chloé Clavel. 2024 · 2024
Closest in time.
Unpacking dpo and ppo: Disentangling best practices for learning from preference feedback
Hamish Ivison, Yizhong Wang, Jiacheng Liu, Zeqiu Wu, Valentina Pyatkin, Nathan Lambert, Noah A Smith, Yejin Choi, and Hannaneh Hajishirzi. 2024 · 2024
Closest in time.
To diverge or not to diverge: A morphosyntactic perspective on machine translation vs human translation
Jiaming Luo, Colin Cherry, and George Foster. 2024 · 2024
Closest in time.
Ai and the problem of knowledge collapse
Andrew J Peterson. 2024 · 2024
Closest in time.
Ai models collapse when trained on recursively generated data
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. 2024 · 2024
Closest in time.
Aya dataset: An open-access collection for multilingual instruction tuning
Shivalika Singh, Freddie Vargus, Daniel D’souza, Börje Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O’Mahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Moura, Dominik Krzemiński, Hakimeh Fadaei, Irem Ergun, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Vu Chien, Sebastian Ruder, Surya Guthikonda, Emad Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet Üstün, Marzieh Fadaee, and Sara Hooker. 2024 · 2024
Closest in time.
Do llamas work in English? on the latent language of multilingual transformers
Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West. 2024 · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. 2024 · 2024
Closest in time.
SafetyBench: Evaluating the safety of large language models
Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2024 · 2024
Closest in time.
From explanations to human-ai co-evolution: charting trajectories towards future user-centric ai
Jürgen Ziegler and Tim Donkers. 2024 · 2024
Closest in time.