Fetching the paper…
Reading the bibliography…
Unlike classical lexical overlap metrics such as BLEU, most current evaluation metrics for machine translation (for example, COMET or BERTScore) are based on black-box large language models.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Trust building with explanation interfaces
Pearl Pu and Li Chen · 2006
Earlier work this paper cites.
Estimating machine translation post-editing effort with HTER
Lucia Specia and Atefeh Farzindar · 2010
Earlier work this paper cites.
Comprehensible classification models: a position paper
Alex A Freitas · 2014
Earlier work this paper cites.
Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics
Arle Lommel, Aljoscha Burchardt, and Hans Uszkoreit · 2014
Earlier work this paper cites.
Adaptive quality estimation for machine translation
Marco Turchi, Antonios Anastasopoulos, José G. C. de Souza, and Matteo Negri · 2014
Earlier work this paper cites.
Accurate evaluation of segment-level machine translation metrics
Yvette Graham, Timothy Baldwin, and Nitika Mathur · 2015
Earlier work this paper cites.
Principles of explanatory debugging to personalize interactive machine learning
Todd Kulesza, Margaret Burnett, Weng-Keen Wong, and Simone Stumpf · 2015
Earlier work this paper cites.
SemEval-2016 task 2: Interpretable semantic textual similarity
Eneko Agirre, Aitor Gonzalez-Agirre, Iñigo Lopez-Gazpio, Montse Maritxalar, German Rigau, and Larraitz Uria · 2016
Earlier work this paper cites.
Can machine translation systems be evaluated by the crowd alone
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel · 2016
Earlier work this paper cites.
Interacting with predictions: Visual inspection of black-box machine learning models
Josua Krause, Adam Perer, and Kenney Ng · 2016
Earlier work this paper cites.
The mythos of model interpretability
Zachary Lipton · 2016
Earlier work this paper cites.
“why should I trust you?”: Explaining the predictions of any classifier
Marco Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Explanation and justification in machine learning : A survey
Or Biran and Courtenay V. Cotton · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Peeking inside the black-box: A survey on explainable artificial intelligence (xai)
Amina Adadi and Mohammed Berrada · 2018
Earlier work this paper cites.
Explainable artificial intelligence: A survey
Filip Karlo Došilović, Mario Brčić, and Nikica Hlupić · 2018
Earlier work this paper cites.
Explaining explanations: An overview of interpretability of machine learning
Leilani H. Gilpin, David Bau, Ben Z. Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal · 2018
Earlier work this paper cites.
A survey of methods for explaining black box models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi · 2018
Earlier work this paper cites.
Teaching categories to human learners with visual explanations
Oisin Mac Aodha, Shihan Su, Yuxin Chen, Pietro Perona, and Yisong Yue · 2018
Earlier work this paper cites.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller · 2018
Earlier work this paper cites.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Earlier work this paper cites.
Explainable artificial intelligence: Understanding, visualizing and interpreting deep learning models
Wojciech Samek, Thomas Wiegand, and Klaus-Robert Müller · 2018
Earlier work this paper cites.
Quality Estimation for other Applications , pages 81–113
Lucia Specia, Carolina Scarton, and Gustavo Henrique Paetzold · 2018
Earlier work this paper cites.
An operation sequence model for explainable neural machine translation
Felix Stahlberg, Danielle Saunders, and Bill Byrne · 2018
Earlier work this paper cites.
A study of reinforcement learning for neural machine translation
Lijun Wu, Fei Tian, Tao Qin, Jianhuang Lai, and Tie-Yan Liu · 2018
Earlier work this paper cites.
One explanation does not fit all: A toolkit and taxonomy of ai explainability techniques
Vijay Arya, Rachel KE Bellamy, Pin-Yu Chen, Amit Dhurandhar, Michael Hind, Samuel C Hoffman, Stephanie Houde, Q Vera Liao, Ronny Luss, Aleksandra Mojsilović, et al · 2019
Earlier work this paper cites.
Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador Garcia, Sergio Gil-Lopez, Daniel Molina, Richard Benjamins, Raja Chatila, and Francisco Herrera · 2019
Earlier work this paper cites.
Visual interaction with deep learning models through collaborative semantic inference
Sebastian Gehrmann, Hendrik Strobelt, Robert Krüger, Hanspeter Pfister, and Alexander M. Rush · 2019
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith · 2019
Earlier work this paper cites.
Studying summarization evaluation metrics in the appropriate scoring range
Maxime Peyrard · 2019
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher · 2019
Earlier work this paper cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 2019
Earlier work this paper cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger · 2019
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao · 2020
Earlier work this paper cites.
Generating hierarchical explanations on text classification via feature interaction detection
Hanjie Chen, Guangtao Zheng, and Yangfeng Ji · 2020
Earlier work this paper cites.
A survey of the state of explainable AI for natural language processing
Marina Danilevsky, Kun Qian, Ranit Aharonov, Yannis Katsis, Ban Kawas, and Prithviraj Sen · 2020
Earlier work this paper cites.
Unsupervised quality estimation for neural machine translation
Marina Fomicheva, Shuo Sun, Lisa Yankovskaya, Frédéric Blain, Francisco Guzmán, Mark Fishel, Nikolaos Aletras, Vishrav Chaudhary, and Lucia Specia · 2020
Earlier work this paper cites.
Explaining black box predictions and unveiling data artifacts through influence functions
Xiaochuang Han, Byron C. Wallace, and Yulia Tsvetkov · 2020
Earlier work this paper cites.
Evaluating explainable AI: Which algorithmic explanations help users predict model behavior?
Peter Hase and Mohit Bansal · 2020
Earlier work this paper cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg · 2020
Earlier work this paper cites.
“why is’ chicago’deceptive?” towards building model-driven tutorials for humans
Vivian Lai, Han Liu, and Chenhao Tan · 2020
Earlier work this paper cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Earlier work this paper cites.
Results of the WMT20 metrics shared task
Nitika Mathur, Johnny Wei, Markus Freitag, Qingsong Ma, and Ondřej Bojar · 2020
Earlier work this paper cites.
TransQuest: Translation quality estimation with cross-lingual transformers
Tharindu Ranasinghe, Constantin Orasan, and Ruslan Mitkov · 2020
Earlier work this paper cites.
TransQuest at WMT2020: Sentence-level direct assessment
Tharindu Ranasinghe, Constantin Orasan, and Ruslan Mitkov · 2020
Earlier work this paper cites.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie · 2020
Earlier work this paper cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Earlier work this paper cites.
An explainable ai decision-support-system to automate loan underwriting
Swati Sachan, Jian-Bo Yang, Dong-Ling Xu, David Eraso Benavides, and Yang Li · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh · 2020
Cited alongside, same era.
Certifai: A common framework to provide explanations and analyse the fairness and robustness of black-box models
Shubham Sharma, Jette Henderson, and Joydeep Ghosh · 2020
Cited alongside, same era.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Brian Thompson and Matt Post · 2020
Cited alongside, same era.
The relationship between trust in ai and trustworthy machine learning technologies
Ehsan Toreini, Mhairi Aitken, Kovila Coopamootoo, Karen Elliott, Carlos Gonzalez Zelaya, and Aad van Moorsel · 2020
Cited alongside, same era.
BERTScore is unfair: On social bias in language model-based metrics for text generation
Tianxiang Sun, Junliang He, Xipeng Qiu, and Xuanjing Huang · 2022
Later among the works it cites.
CrossQE: HW-TSC 2022 submission for the quality estimation shared task
Shimin Tao, Su Chang, Ma Miaomiao, Hao Yang, Xiang Geng, Shujian Huang, Min Zhang, Jiaxin Guo, Minghan Wang, and Yinglu Li · 2022
Later among the works it cites.
Layer or representation space: What makes BERT-based evaluation metrics robust?
Doan Nam Long Vu, Nafise Sadat Moosavi, and Steffen Eger · 2022
Later among the works it cites.
A fine-grained interpretability evaluation benchmark for neural NLP
Lijie Wang, Yaozong Shen, Shuyuan Peng, Shuai Zhang, Xinyan Xiao, Hao Liu, Hongxuan Tang, Ying Chen, Hua Wu, and Haifeng Wang · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Effect of confidence and explanation on accuracy and trust calibration in ai-assisted decision making
Yunfeng Zhang, Q Vera Liao, and Rachel KE Bellamy · 2020
Cited alongside, same era.
On the limitations of cross-lingual encoders as exposed by reference-free machine translation evaluation
Wei Zhao, Goran Glavaš, Maxime Peyrard, Yang Gao, Robert West, and Steffen Eger · 2020
Cited alongside, same era.
Explaining neural network predictions on sentence pairs via learning word-group masks
Hanjie Chen, Song Feng, Jatin Ganhotra, Hui Wan, Chulaka Gunasekara, Sachindra Joshi, and Yangfeng Ji · 2021
Cited alongside, same era.
Semi-automated data labeling
Michael Desmond, Evelyn Duesterwald, Kristina Brimijoin, Michelle Brachman, and Qian Pan · 2021
Cited alongside, same era.
Explaining errors in machine translation with absolute gradient ensembles
Melda Eksi, Erik Gelbing, Jonathan Stieber, and Chi Viet Vu · 2021
Cited alongside, same era.
The Eval4NLP shared task on explainable quality estimation: Overview and results
Marina Fomicheva, Piyawat Lertvittayakumjorn, Wei Zhao, Steffen Eger, and Yang Gao · 2021
Cited alongside, same era.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey · 2021
Cited alongside, same era.
Findings of the wmt 2022 shared task on quality estimation
Chrysoula Zerva, Frédéric Blain, Ricardo Rei and Piyawat Lertvittayakumjorn, José G. C. de Souza, Steffen Eger, Diptesh Kanojia, Duarte Alves, Constantin Orăsan, Marina Fomicheva, André F. T. Martins, and Lucia Specia · 2022
Later among the works it cites.
On the explainability of natural language processing deep models
Julia El Zini and Mariette Awad · 2022
Later among the works it cites.
ACES: Translation accuracy challenge sets at WMT 2023
Chantal Amrhein, Nikita Moghe, and Liane Guillou · 2023
Closest in time.
Palm 2 technical report
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Tachard Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Z. Chen, Eric Chu, J. Clark, Laurent El Shafey, Yanping Huang, Kathleen S. Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernández Abrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan A. Botha, James Bradbury, Siddhartha Brahma, Kevin Michael Brooks, Michele Catasta, Yongzhou Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, C Crépy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, M. C. D’iaz, Nan Du, Ethan Dyer, Vladimir Feinberg, Fan Feng, Vlad Fienber, Markus Freitag, Xavier García, Sebastian Gehrmann, Lucas González, Guy Gur-Ari, Steven Hand, Hadi Hashemi, Le Hou, Joshua Howland, An Ren Hu, Jeffrey Hui, Jeremy Hurwitz, Michael Isard, Abe Ittycheriah, Matthew Jagielski, Wen Hao Jia, Kathleen Kenealy, Maxim Krikun, Sneha Kudugunta, Chang Lan, Katherine Lee, Benjamin Lee, Eric Li, Mu-Li Li, Wei Li, Yaguang Li, Jun Yu Li, Hyeontaek Lim, Han Lin, Zhong-Zhong Liu, Frederick Liu, Marcello Maggioni, Aroma Mahendru, Joshua Maynez, Vedant Misra, Maysam Moussalem, Zachary Nado, John Nham, Eric Ni, Andrew Nystrom, Alicia Parrish, Marie Pellat, Martin Polacek, Oleksandr Polozov, Reiner Pope, Siyuan Qiao, Emily Reif, Bryan Richter, Parker Riley, Alexandra Ros, Aurko Roy, Brennan Saeta, Rajkumar Samuel, Renee Marie Shelby, Ambrose Slone, Daniel Smilkov, David R. So, Daniela Sohn, Simon Tokumine, Dasha Valter, Vijay Vasudevan, Kiran Vodrahalli, Xuezhi Wang, Pidong Wang, Zirui Wang, Tao Wang, John Wieting, Yuhuai Wu, Ke Xu, Yunhan Xu, Lin Wu Xue, Pengcheng Yin, Jiahui Yu, Qiaoling Zhang, Steven Zheng, Ce Zheng, Wei Zhou, Denny Zhou, Slav Petrov, and Yonghui Wu · 2023
Closest in time.
Challenging the state-of-the-art machine translation metrics from a linguistic perspective
Eleftherios Avramidis, Shushen Manakhimova, Vivien Macketanz, and Sebastian Möller · 2023
Closest in time.
Towards fine-grained information: Identifying the type and location of translation errors
Keqin Bao, Yu Wan, Dayiheng Liu, Baosong Yang, Wenqiang Lei, Xiangnan He, Derek F.Wong, and Jun Xie · 2023
Closest in time.
UScore: An effective approach to fully unsupervised evaluation metrics for machine translation
Jonas Belouadi and Steffen Eger · 2023
Closest in time.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Closest in time.
Findings of the WMT 2023 shared task on quality estimation
Frederic Blain, Chrysoula Zerva, Ricardo Ribeiro, Nuno M. Guerreiro, Diptesh Kanojia, José G. C. de Souza, Beatriz Silva, Tânia Vaz, Yan Jingxuan, Fatemeh Azadi, Constantin Orasan, and André Martins · 2023
Closest in time.
Benchmarking and survey of explanation methods for black box models
Francesco Bodria, Fosca Giannotti, Riccardo Guidotti, Francesca Naretto, Dino Pedreschi, and Salvatore Rinzivillo · 2023
Closest in time.
MENLI: Robust evaluation metrics from natural language inference
Yanran Chen and Steffen Eger · 2023
Closest in time.
Exploring the use of large language models for reference-free text quality evaluation: An empirical study
Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu · 2023
Closest in time.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee · 2023
Closest in time.
The devil is in the errors: Leveraging large language models for fine-grained machine translation evaluation
Patrick Fernandes, Daniel Deutsch, Mara Finkelstein, Parker Riley, André Martins, Graham Neubig, Ankush Garg, Jonathan Clark, Markus Freitag, and Orhan Firat · 2023
Closest in time.
Dear xai community, we need to talk!
Timo Freiesleben and Gunnar König · 2023
Closest in time.
Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent
Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster · 2023
Closest in time.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Closest in time.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam · 2023
Closest in time.
Unify word-level and span-level tasks: NJUNLP’s participation for the WMT2023 quality estimation shared task
Xiang Geng, Zhejian Lai, Yu Zhang, Shimin Tao, Hao Yang, Jiajun Chen, and Shujian Huang · 2023
Closest in time.
ROSCOE: A suite of metrics for scoring step-by-step reasoning
Olga Golovneva, Moya Peng Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz · 2023
Closest in time.
A survey of adversarial defences and robustness in nlp
Shreya Goyal, Sumanth Doddapaneni, Mitesh M. Khapra, and Balaraman Ravindran · 2023
Closest in time.
Looking for a needle in a haystack: A comprehensive study of hallucinations in neural machine translation
Nuno M. Guerreiro, Elena Voita, and André Martins · 2023
Closest in time.
Rationalization for explainable nlp: A survey
Sai Gurrapu, Ajay Kulkarni, Lifu Huang, Ismini Lourentzou, Laura J. Freeman, and Feras A. Batarseh · 2023
Closest in time.
On the blind spots of model-based evaluation metrics for text generation
Tianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James Glass, and Yulia Tsvetkov · 2023
Closest in time.
Diagnosing ai explanation methods with folk concepts of behavior
Alon Jacovi, Jasmijn Bastings, Sebastian Gehrmann, Yoav Goldberg, and Katja Filippova · 2023
Closest in time.
Exploring chatgpt’s ability to rank content: A preliminary study on consistency with human preferences, 2023
Yunjie Ji, Yan Gong, Yiping Peng, Chao Ni, Peiyan Sun, Dongyu Pan, Baochang Ma, and Xiangang Li · 2023
Closest in time.
Rethinking ai explainability and plausibility
Weina Jin, Xiaoxiao Li, and Ghassan Hamarneh · 2023
Closest in time.
DATScore: Evaluating translation with data augmented translations
Moussa Kamal Eddine, Guokan Shang, and Michalis Vazirgiannis · 2023
Closest in time.
GEMBA-MQM: Detecting translation quality error spans with GPT-4
Tom Kocmi and Christian Federmann · 2023
Closest in time.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann · 2023
Closest in time.
The eval4nlp 2023 shared task on prompting large language models as explainable metrics
Christoph Leiter, Juri Opitz, Daniel Deutsch, Yang Gao, Rotem Dror, and Steffen Eger · 2023
Closest in time.
HW-TSC 2023 submission for the quality estimation shared task
Yuang Li, Chang Su, Ming Zhu, Mengyao Piao, Xinglin Lyu, Min Zhang, and Hao Yang · 2023
Closest in time.
G-eval: Nlg evaluation using gpt-4 with better human alignment, 2023
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu · 2023
Closest in time.
Metric score landscape challenge (MSLC23): Understanding metrics’ performance on a wider landscape of translation quality
Chi-kiu Lo, Samuel Larkin, and Rebecca Knowles · 2023
Closest in time.
Toward human-like evaluation for natural language generation with error analysis
Qingyu Lu, Liang Ding, Liping Xie, Kanjian Zhang, Derek F. Wong, and Dacheng Tao · 2023
Closest in time.
Introducing chatgpt
OpenAI · 2023
Closest in time.
There’s no data like better data: Using QE metrics for MT data filtering
Jan-Thorsten Peter, David Vilar, Daniel Deutsch, Mara Finkelstein, Juraj Juraska, and Markus Freitag · 2023
Closest in time.
Aligning neural machine translation models: Human feedback in training and inference
Miguel Moura Ramos, Patrick Fernandes, António Farinhas, and Andr’e F. T. Martins · 2023
Closest in time.
Scaling up CometKiwi: Unbabel-IST 2023 submission for the quality estimation shared task
Ricardo Rei, Nuno M. Guerreiro, José Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, José G. C. de Souza, and André Martins · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Interpretability in activation space analysis of transformers: A focused survey
Soniya Vijayakumar · 2023
Closest in time.
Is ChatGPT a good NLG evaluator? a preliminary study
Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou · 2023
Closest in time.
INSTRUCTSCORE: Towards explainable text generation evaluation with automatic feedback
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Wang, and Lei Li · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Closest in time.
DiscoScore: Evaluating text generation with BERT and discourse coherence
Wei Zhao, Michael Strube, and Steffen Eger · 2023
Closest in time.
Poor man’s quality estimation: Predicting reference-based MT metrics without the reference
Vilém Zouhar, Shehzaad Dhuliawala, Wangchunshu Zhou, Nico Daheim, Tom Kocmi, Yuchen Eleanor Jiang, and Mrinmaya Sachan · 2023
Closest in time.
SemEval-2015 task 2: Semantic textual similarity, English, Spanish and pilot on interpretability
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Inigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, German Rigau, Larraitz Uria, and Janyce Wiebe · 2045
Closest in time.
Machine learning interpretability: A survey on methods and metrics
Diogo V. Carvalho, Eduardo M. Pereira, and Jaime S. Cardoso · 2079
Closest in time.