Fetching the paper…
Reading the bibliography…
We present a comprehensive evaluation of large language models for multilingual readability assessment.
Subjective assessment of text complexity: A dataset for German language
Babak Naderi, Salar Mohtaj, Kaspar Ensikat, and Sebastian Möller. 2019 · 1904
Earlier work this paper cites.
Adaptation of deep bidirectional multilingual transformers for russian language
Yuri Kuratov and Mikhail Arkhipov. 2019 · 1905
Earlier work this paper cites.
Text readability assessment for second language learners
Menglin Xia, Ekaterina Kochmar, and Ted Briscoe. 2019 · 1906
Earlier work this paper cites.
Automated readability index , volume 66
Edgar A Smith and RJ Senter. 1967 · 1967
Earlier work this paper cites.
Derivation of new readability formulas (Automated Readability Index, Fog Count and Flesch Reading Ease Formula) for navy enlisted personnel
J Peter Kincaid, Robert P. Fishburne Jr., Richard L. Rogers, and Brad S. Chissom. 1975 · 1975
Earlier work this paper cites.
The measurement of personal values in survey research: A test of alternative rating procedures
John A McCarty and Larry J Shrum. 2000 · 2000
Earlier work this paper cites.
Arabert: Transformer-based model for arabic language understanding
Wissam Antoun, Fady Baly, and Hazem Hajj. 2020 · 2003
Earlier work this paper cites.
Inter-coder agreement for computational linguistics
Ron Artstein and Massimo Poesio. 2008 · 2008
Earlier work this paper cites.
2010 i2b2/va challenge on concepts, assertions, and relations in clinical text
Özlem Uzuner, Brett R South, Shuying Shen, and Scott L DuVall. 2011 · 2010
Earlier work this paper cites.
On improving the accuracy of readability classification using insights from second language acquisition
Sowmya Vajjala and Detmar Meurers. 2012 · 2012
Earlier work this paper cites.
LABR: A large scale Arabic book reviews dataset
Mohamed Aly and Amir Atiya. 2013 · 2013
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
P. Malo, A. Sinha, P. Korhonen, J. Wallenius, and P. Takala. 2014 · 2014
Earlier work this paper cites.
A corpus and grammatical browsing system for remedial EFL learners
Kiyomi Chujo, Kathryn Oghigian, and Shiro Akasegawa. 2015 · 2015
Earlier work this paper cites.
Building large Arabic multi-domain resources for sentiment analysis
Hady ElSahar and Samhaa R El-Beltagy. 2015 · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Problems in current text simplification research: New data can help
Wei Xu, Chris Callison-Burch, and Courtney Napoles. 2015 · 2015
Earlier work this paper cites.
Aspect based sentiment analysis in Hindi: resource creation and evaluation
Md Shad Akhtar, Asif Ekbal, and Pushpak Bhattacharyya. 2016 · 2016
Earlier work this paper cites.
All mixed up? Finding the optimal feature set for general readability prediction and its application to English and Dutch
Orphée De Clercq and Véronique Hoste. 2016 · 2016
Earlier work this paper cites.
OSMAN — a novel Arabic readability metric
Mahmoud El-Haj and Paul Rayson. 2016 · 2016
Earlier work this paper cites.
OpenSubtitles2016: Extracting large parallel corpora from movie and tv subtitles
Pierre Lison and Jörg Tiedemann. 2016 · 2016
Earlier work this paper cites.
SemEval-2016 task 3: Community question answering
Preslav Nakov, Lluís Màrquez, Alessandro Moschitti, Walid Magdy, Hamdy Mubarak, Abed Alhakim Freihat, Jim Glass, and Bilal Randeree. 2016 · 2016
Earlier work this paper cites.
What to do about non-standard (or non-canonical) language in NLP
Barbara Plank. 2016 · 2016
Earlier work this paper cites.
Text readability assessment for second language learners
Menglin Xia, Ekaterina Kochmar, and Ted Briscoe. 2016 · 2016
Earlier work this paper cites.
The United Nations Parallel Corpus v1. 0
Michał Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016 · 2016
Earlier work this paper cites.
DailyDialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017 · 2017
Earlier work this paper cites.
Quora question pairs
Quora.com. 2017 · 2017
Earlier work this paper cites.
Automatic assessment of absolute sentence complexity
Sanja Štajner, Simone Paolo Ponzetto, and Heiner Stuckenschmidt. 2017 · 2017
Earlier work this paper cites.
Is this sentence difficult? do you agree?
Dominique Brunato, Lorenzo De Mattei, Felice Dell’Orletta, Benedetta Iavarone, and Giulia Venturi. 2018 · 2018
Earlier work this paper cites.
Assessing the readability of web search results for searchers with dyslexia
Adam Fourney, Meredith Ringel Morris, Abdullah Ali, and Laura Vonessen. 2018 · 2018
Earlier work this paper cites.
Decoupling strategy and generation in negotiation dialogues
He He, Derek Chen, Anusha Balakrishnan, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Joint learning of frequency and word embeddings for multilingual readability assessment
Dieu-Thu Le, Cam-Tu Nguyen, and Xiaoliang Wang. 2018 · 2018
Earlier work this paper cites.
A neural local coherence model for text quality assessment
Mohsen Mesgar and Michael Strube. 2018 · 2018
Earlier work this paper cites.
A dataset and reranking method for multimodal mt of user-generated image captions
Shigehiko Schamoni, Julian Hitschler, and Stefan Riezler. 2018 · 2018
Earlier work this paper cites.
OneStopEnglish corpus: A new corpus for automatic readability assessment and text simplification
Sowmya Vajjala and Ivana Lučić. 2018 · 2018
Cited alongside, same era.
Controlling text complexity in neural machine translation
Sweta Agrawal and Marine Carpuat. 2019 · 2019
Cited alongside, same era.
Multiattentive recurrent neural network architecture for multilingual readability assessment
Ion Madrazo Azpiazu and Maria Soledad Pera. 2019 · 2019
Cited alongside, same era.
CoFiF: A corpus of financial reports in french language
Tobias Daudert and Sina Ahmadi. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Sentiment analysis of product reviews in russian using convolutional neural networks
Empathetic BERT2BERT conversational model: Learning Arabic language generation with little data
Tarek Naous, Wissam Antoun, Reem Mahmoud, and Hazem Hajj. 2021 · 2021
Later among the works it cites.
Cross-lingual leveled reading based on language-invariant features
Simin Rao, Hua Zheng, and Sujian Li. 2021 · 2021
Later among the works it cites.
An ensemble-based hotel recommender system using sentiment analysis and aspect categorization of hotel reviews
Biswarup Ray, Avishek Garain, and Ram Sarkar. 2021 · 2021
Later among the works it cites.
From masked language modeling to translation: Non-English auxiliary tasks improve zero-shot spoken language understanding
Rob van der Goot, Ibrahim Sharaf, Aizhan Imankulova, Ahmet Üstün, Marija Stepanović, Alan Ramponi, Siti Oryza Khairunnisa, Mamoru Komachi, and Barbara Plank. 2021 · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergey Smetanin and Michail Komarov. 2019 · 2019
Cited alongside, same era.
Fine-grained spoiler detection from large-scale review corpora
Mengting Wan, Rishabh Misra, Ndapandula Nakashole, and Julian McAuley. 2019 · 2019
Cited alongside, same era.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. 2019 · 2019
Cited alongside, same era.
Humor detection: A transformer gets the last laugh
Orion Weller and Kevin Seppi. 2019 · 2019
Cited alongside, same era.
Cross-lingual transfer learning for intent detection of covid-19 utterances
Abhinav Arora, Akshat Shrivastava, Mrinal Mohit, Lorena Sainz-Maza Lecanda, and Ahmed Aly. 2020 · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
FQuAD: French question answering dataset
Martin d’Hoffschmidt, Wacim Belblidia, Quentin Heinrich, Tom Brendlé, and Maxime Vidal. 2020 · 2020
Cited alongside, same era.
A dataset for detecting humor in Arabic text
Hend Al-Khalifa, Fetoun AlZahrani, Hala Qawara, Reema AlRowais, Sawsan Alowa, and Luluh AlDhubayi. 2022 · 2022
Later among the works it cites.
A novel methodology for Arabic news classification
Marco Alfonse and Mariam Gawich. 2022 · 2022
Later among the works it cites.
CEFR-based sentence difficulty annotation and assessment
Yuki Arase, Satoru Uchida, and Tomoyuki Kajiwara. 2022 · 2022
Later among the works it cites.
Automatic readability assessment of German sentences with transformer ensembles
Patrick Gustav Blaneck, Tobias Bornheim, Niklas Grieger, and Stephan Bialonski. 2022 · 2022
Later among the works it cites.
AraT5: Text-to-text transformers for arabic language generation
AbdelRahim Elmadany, Muhammad Abdul-Mageed, et al. 2022 · 2022
Later among the works it cites.
ZAEBUC: An annotated arabic-english bilingual writer corpus
Nizar Habash and David Palfreyman. 2022 · 2022
Later among the works it cites.
A baseline readability model for Cebuano
Joseph Marvin Imperial, Lloyd Lois Antonie Reyes, Michael Antonio Ibanez, Ranz Sapinit, and Mohammed Hussien. 2022 · 2022
Later among the works it cites.
HLDC: Hindi legal documents corpus
Arnav Kapoor, Mudit Dhawan, Anmol Goel, TH Arjun, Akshala Bhatnagar, Vibhu Agrawal, Amul Agrawal, Arnab Bhattacharya, Ponnurangam Kumaraguru, and Ashutosh Modi. 2022 · 2022
Later among the works it cites.
A neural pairwise ranking model for readability assessment
Justin Lee and Sowmya Vajjala. 2022 · 2022
Later among the works it cites.
Rishabh Misra. 2022 · 2022
Later among the works it cites.
Attention based video captioning framework for Hindi
Alok Singh, Thoudam Doren Singh, and Sivaji Bandyopadhyay. 2022 · 2022
Later among the works it cites.
Rusentitweet: A sentiment analysis dataset of general domain tweets in russian
Sergey Smetanin. 2022 · 2022
Later among the works it cites.
Trends, limitations and open challenges in automatic readability assessment research
Sowmya Vajjala. 2022 · 2022
Later among the works it cites.
MDIA: A benchmark for multilingual dialogue generation in 46 languages
Qingyu Zhang, Xiaoyu Shen, Ernie Chang, Jidong Ge, and Pengke Chen. 2022 · 2022
Later among the works it cites.
Stanceosaurus: Classifying stance towards multilingual misinformation
Jonathan Zheng, Ashutosh Baheti, Tarek Naous, Wei Xu, and Alan Ritter. 2022 · 2022
Later among the works it cites.
Learning to paraphrase sentences to different complexity levels
Alison Chi, Li-Kuang Chen, Yi-Chen Chang, Shu-Hui Lee, and Jason S Chang. 2023 · 2023
Closest in time.
Simplicity level estimate (sle): A learned reference-less metric for sentence simplification
Liam Cripwell, Joël Legrand, and Claire Gardent. 2023 · 2023
Closest in time.
Hindi movie reviews dataset
HindiMovieReviews · 2023
Closest in time.
Automatic readability assessment for closely related languages
Joseph Marvin Imperial and Ekaterina Kochmar. 2023 · 2023
Closest in time.
LENS: A learnable evaluation metric for text simplification
Mounica Maddela, Yao Dou, David Heineman, and Wei Xu. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
MultiCapCLIP: Auto-encoding prompts for zero-shot multilingual visual captioning
Bang Yang, Fenglin Liu, Xian Wu, Yaowei Wang, Xu Sun, and Yuexian Zou. 2023 · 2023
Closest in time.
The SAMER arabic text simplification corpus
Bashar Alhafni, Reem Hazim, Juan David Pineros Liberato, Muhamed Al Khalil, and Nizar Habash. 2024 · 2024
Closest in time.
Aya 23: Open weight releases to further multilingual progress
Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Kelly Marchisio, Sebastian Ruder, et al. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
Aya dataset: An open-access collection for multilingual instruction tuning
Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, et al. 2024 · 2024
Closest in time.
Aya model: An instruction finetuned open-access multilingual language model
Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, et al. 2024 · 2024
Closest in time.