Fetching the paper…
Reading the bibliography…
While general-purpose large language models (LLMs) demonstrate proficiency on multiple tasks within the domain of translation, approaches based on open LLMs are competitive only when specializing on a single task.
Wikimatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán · 1907
Earlier work this paper cites.
Ccmatrix: Mining billions of high-quality parallel sentences on the web
Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, and Armand Joulin · 1911
Earlier work this paper cites.
Ccnet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave · 1911
Earlier work this paper cites.
Speech understanding systems: A summary of results of the five-year research effort at carnegie mellon university., 1977
Raj Reddy · 1977
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn · 2004
Earlier work this paper cites.
Improving massively multilingual neural machine translation and zero-shot translation
Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich · 2004
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn · 2005
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Rich Schwartz, Linnea Micciulla, and John Makhoul · 2006
Earlier work this paper cites.
MultiUN: A multilingual corpus from united nation documents
Andreas Eisele and Yu Chen · 2010
Earlier work this paper cites.
KenLM: Faster and smaller language model queries
Kenneth Heafield · 2011
Earlier work this paper cites.
Parallel data, tools and interfaces in opus
Jörg Tiedemann · 2012
Earlier work this paper cites.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel · 2013
Earlier work this paper cites.
Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics
Arle Lommel, Aljoscha Burchardt, and Hans Uszkoreit · 2014
Earlier work this paper cites.
Creating a massively parallel Bible corpus
Thomas Mayer and Michael Cysouw · 2014
Earlier work this paper cites.
The CoNLL-2014 shared task on grammatical error correction
Hwee Tou Ng, Siew Mei Wu, Ted Briscoe, Christian Hadiwinoto, Raymond Hendy Susanto, and Christopher Bryant · 2014
Earlier work this paper cites.
Building subject-aligned comparable corpora and mining it for truly parallel sentence pairs
Krzysztof Wołk and Krzysztof Marasek · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović · 2015
Earlier work this paper cites.
Automatic extraction of learner errors in ESL sentences using linguistically enhanced alignments
Mariano Felice, Christopher Bryant, and Ted Briscoe · 2016
Earlier work this paper cites.
The United Nations parallel corpus v1.0
Michał Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen · 2016
Earlier work this paper cites.
Automatic annotation and evaluation of error types for grammatical error correction
Christopher Bryant, Mariano Felice, and Ted Briscoe · 2017
Earlier work this paper cites.
Tilde MODEL - multilingual open data for EU languages
Roberts Rozis and Raivis Skadiņš · 2017
Earlier work this paper cites.
Translation quality and productivity: A study on rich morphology languages
Lucia Specia, Kim Harris, Frédéric Blain, Aljoscha Burchardt, Viviven Macketanz, Inguna Skadin, Matteo Negri, and Marco Turchi · 2017
Earlier work this paper cites.
A multilayer convolutional encoder-decoder neural network for grammatical error correction
Shamil Chollampatt and Hwee Tou Ng · 2018
Earlier work this paper cites.
Prompsit’s submission to WMT 2018 parallel corpus filtering shared task
Víctor M. Sánchez-Cartagena, Marta Bañón, Sergio Ortiz-Rojas, and Gema Ramírez · 2018
Earlier work this paper cites.
A large parallel corpus of full-text scientific articles
Felipe Soares, Viviane Moreira, and Karin Becker · 2018
Earlier work this paper cites.
ParaCrawl: Web-scale parallel corpora for the languages of the EU
Miquel Esplà, Mikel Forcada, Gema Ramírez-Sánchez, and Hieu Hoang · 2019
Earlier work this paper cites.
PAWS-X: A cross-lingual adversarial dataset for paraphrase identification
Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge · 2019
Earlier work this paper cites.
The effect of translationese in machine translation test sets
Mike Zhang and Antonio Toral · 2019
Earlier work this paper cites.
TICO-19: the translation initiative for COvid-19
Antonios Anastasopoulos, Alessandro Cattelan, Zi-Yi Dou, Marcello Federico, Christian Federmann, Dmitriy Genzel, Franscisco Guzmán, Junjie Hu, Macduff Hughes, Philipp Koehn, Rosie Lazar, Will Lewis, Graham Neubig, Mengmeng Niu, Alp Öktem, Eric Paquin, Grace Tang, and Sylwia Tur · 2020
Earlier work this paper cites.
Is MAP decoding all you need? the inadequacy of the mode in neural machine translation
Bryan Eikema and Wilker Aziz · 2020
Earlier work this paper cites.
CCAligned: A massive collection of cross-lingual web-document pairs
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, and Philipp Koehn · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
Bifixer and bicleaner: two open-source tools to clean your parallel data
Gema Ramírez-Sánchez, Jaume Zaragoza-Bernabeu, Marta Bañón, and Sergio Ortiz Rojas · 2020
Cited alongside, same era.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
Translationese as a language in “multilingual” NMT
Parker Riley, Isaac Caswell, Markus Freitag, and David Grangier · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
The devil is in the errors: Leveraging large language models for fine-grained machine translation evaluation
Patrick Fernandes, Daniel Deutsch, Mara Finkelstein, Parker Riley, André Martins, Graham Neubig, Ankush Garg, Jonathan Clark, Markus Freitag, and Orhan Firat · 2023
Later among the works it cites.
MultiCoNER v2: a large multilingual dataset for fine-grained and noisy named entity recognition
Besnik Fetahu, Zhiyu Chen, Sudipta Kar, Oleg Rokhlenko, and Shervin Malmasi · 2023
Later among the works it cites.
Results of wmt23 metrics shared task: Metrics might be guilty but references are not innocent
Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster · 2023
Later among the works it cites.
xCOMET: Transparent machine translation evaluation through fine-grained error detection
Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André F. T. Martins · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thibault Sellam, Dipanjan Das, and Ankur Parikh · 2020
Cited alongside, same era.
The Tatoeba Translation Challenge – Realistic data sets for low resource and multilingual MT
Jörg Tiedemann · 2020
Cited alongside, same era.
Cows-l2h: A corpus of spanish learner writing
Aaron Yamada, Sam Davidson, Paloma Fernández-Mira, Agustina Carando, Kenji Sagae, and Claudia Sánchez-Gutiérrez · 2020
Cited alongside, same era.
Philip Williams and Barry Haddow · 2021
Cited alongside, same era.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel · 2021
Cited alongside, same era.
What are the best systems? new perspectives on nlp benchmarking
Pierre Colombo, Nathan Noiry, Ekhine Irurozki, and Stéphan Clémençon · 2022
Cited alongside, same era.
Anna Currey, Maria Nadejde, Raghavendra Pappagari, Mia Mayer, Stanislas Lauly, Xing Niu, Benjamin Hsu, and Georgiana Dinu · 2022
Cited alongside, same era.
Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla · 2023
Later among the works it cites.
Can LLMs learn from a single example?, howpublished = https://www.fast.ai/posts/2023-09-04-learning-jumps/ , note = Accessed: 2024-02-22, 2023
Jeremy Howard and Jonathan Whitaker · 2023
Later among the works it cites.
Towards effective disambiguation for machine translation with large language models
Vivek Iyer, Pinzhen Chen, and Alexandra Birch · 2023
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
GEMBA-MQM: Detecting translation quality error spans with GPT-4
Tom Kocmi and Christian Federmann · 2023
Later among the works it cites.
Findings of the 2023 conference on machine translation (WMT23): LLMs are here but not quite there yet
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Philipp Koehn, Benjamin Marie, Christof Monz, Makoto Morishita, Kenton Murray, Makoto Nagata, Toshiaki Nakazawa, Martin Popel, Maja Popović, and Mariya Shmatova · 2023
Later among the works it cites.
Chipnemo: Domain-adapted llms for chip design
Mingjie Liu, Teodor-Dumitru Ene, Robert Kirby, Chris Cheng, Nathaniel Pinckney, Rongjian Liang, Jonah Alben, Himyanshu Anand, Sanmitra Banerjee, Ismet Bayraktaroglu, Bonita Bhaskaran, Bryan Catanzaro, Arjun Chaudhuri, Sharon Clay, Bill Dally, Laura Dang, Parikshit Deshpande, Siddhanth Dhodhi, Sameer Halepete, Eric Hill, Jiashang Hu, Sumit Jain, Brucek Khailany, George Kokai, Kishor Kunal, Xiaowei Li, Charley Lind, Hao Liu, Stuart Oberman, Sujeet Omar, Sreedhar Pratty, Jonathan Raiman, Ambar Sarkar, Zhengjiang Shao, Hanfei Sun, Pratik P Suthar, Varun Tej, Walker Turner, Kaizhe Xu, and Haoxing Ren · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, and Adam Roberts · 2023
Later among the works it cites.
Supervised fine-tuning and direct preference optimization on intel gaudi2
Kaokao Lv, Wenxin Zhang, and Haihao Shen · 2023
Later among the works it cites.
Findings of the WMT 2023 biomedical translation shared task: Evaluation of ChatGPT 3.5 as a comparison system
Mariana Neves, Antonio Jimeno Yepes, Aurélie Névéol, Rachel Bawden, Giorgio Maria Di Nunzio, Roland Roller, Philippe Thomas, Federica Vezzani, Maika Vicente Navarro, Lana Yeganova, Dina Wiemann, and Cristian Grozea · 2023
Later among the works it cites.
URL https://github.com/openai/openai-python/blob/release-v0.28.1/chatml.md
Open AI, 2023 · 2023
Later among the works it cites.
Sabiá: Portuguese large language models
Ramon Pires, Hugo Abonizio, Thales Sales Almeida, and Rodrigo Nogueira · 2023
Later among the works it cites.
Leveraging GPT-4 for automatic translation post-editing
Vikas Raunak, Amr Sharaf, Yiren Wang, Hany Awadalla, and Arul Menezes · 2023
Later among the works it cites.
Scaling up CometKiwi: Unbabel-IST 2023 submission for the quality estimation shared task
Ricardo Rei, Nuno M. Guerreiro, José Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, José G. C. de Souza, and André Martins · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, and Thomas Wolf · 2023
Later among the works it cites.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
PolyLM: An Open Source Polyglot Large Language Model
Xiangpeng Wei, Haoran Wei, Huan Lin, Tianhao Li, Pei Zhang, Xingzhang Ren, Mei Li, Yu Wan, Zhiwei Cao, Binbin Xie, Tianxiang Hu, Shangjie Li, Binyuan Hui, Bowen Yu, Dayiheng Liu, Baosong Yang, Fei Huang, and Jun Xie · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann · 2023
Later among the works it cites.
INSTRUCTSCORE: Towards explainable text generation evaluation with automatic feedback
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Wang, and Lei Li · 2023
Later among the works it cites.
LIMA: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, LILI YU, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy · 2023
Later among the works it cites.
Croissantllm: A truly bilingual french-english language model
Manuel Faysse, Patrick Fernandes, Nuno M. Guerreiro, António Loison, Duarte M. Alves, Caio Corro, Nicolas Boizard, João Alves, Ricardo Rei, Pedro H. Martins, Antoni Bigata Casademunt, François Yvon, André F. T. Martins, Gautier Viaud, Céline Hudelot, and Pierre Colombo · 2024
Closest in time.
Gemma: Open Models Based on Gemini Research and Technology, howpublished = https://blog.google/technology/developers/gemma-open-models/ , note = Accessed: 2024-02-27, 2024
Google DeepMind Gemma Team · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Closest in time.
Navigating the metrics maze: Reconciling score magnitudes and accuracies
Tom Kocmi, Vilém Zouhar, Christian Federmann, and Matt Post · 2024
Closest in time.
MAmmoTH: Building math generalist models through hybrid instruction tuning
Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen · 2024
Closest in time.
Long is more for alignment: A simple but tough-to-beat baseline for instruction fine-tuning
Hao Zhao, Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Closest in time.