Fetching the paper…
Reading the bibliography…
In the evolving NLP landscape, benchmarks serve as yardsticks for gauging progress.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Contractnli: A dataset for document-level natural language inference for contracts
Yuta Koreeda and Christopher D Manning. 2021 · 1919
Earlier work this paper cites.
The trec-8 question answering track evaluation
Ellen M Voorhees, Dawn M Tice, et al. 1999 · 1999
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Extracting contract elements
Ilias Chalkidis, Ion Androutsopoulos, and Achilleas Michos. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin. 2018 · 2018
Earlier work this paper cites.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi. 2018 · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Earlier work this paper cites.
Lener-br: a dataset for named entity recognition in brazilian legal text
Pedro Henrique Luz de Araujo, Teófilo E de Campos, Renato RR de Oliveira, Matheus Stauffer, Samuel Couto, and Paulo Bermejo. 2018 · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. 2018 · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
A comparative study of summarization algorithms applied to legal case judgments
Paheli Bhattacharya, Kaustubh Hiware, Subham Rajgaria, Nilay Pochhi, Kripabandhu Ghosh, and Saptarshi Ghosh. 2019 · 2019
Earlier work this paper cites.
Large-scale multi-label text classification on eu legislation
Ilias Chalkidis, Emmanouil Fergadiotis, Prodromos Malakasiotis, and Ion Androutsopoulos. 2019 · 2019
Earlier work this paper cites.
Billsum: A corpus for automatic summarization of us legislation
Anastassia Kornilova and Vladimir Eidelman. 2019 · 2019
Earlier work this paper cites.
Claudette: an automated detector of potentially unfair clauses in online terms of service
Marco Lippi, Przemysław Pałka, Giuseppe Contissa, Francesca Lagioia, Hans-Wolfgang Micklitz, Giovanni Sartor, and Paolo Torroni. 2019 · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
Question answering for privacy policies: Combining computational and legal perspectives
Abhilasha Ravichander and W Alan. 2019 · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019 · 2019
Earlier work this paper cites.
Automatic summarization of legal decisions using iterative masking of predictive sentences
Linwu Zhong, Ziyi Zhong, Zinian Zhao, Siyuan Wang, Kevin D Ashley, and Matthias Grabmair. 2019 · 2019
Earlier work this paper cites.
Etc: Encoding long and structured inputs in transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020 · 2020
Earlier work this paper cites.
Legal-bert: The muppets straight out of law school
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos. 2020 · 2020
Earlier work this paper cites.
Joint entity and relation extraction for legal documents with legal feature enhancement
Yanguang Chen, Yuanyuan Sun, Zhihao Yang, and Hongfei Lin. 2020 · 2020
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020 · 2020
Earlier work this paper cites.
Answering legal questions by learning neural attentive text representation
Phi Manh Kien, Ha-Thanh Nguyen, Ngo Xuan Bach, Vu Tran, Minh Le Nguyen, and Tu Minh Phuong. 2020 · 2020
Earlier work this paper cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Earlier work this paper cites.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
Bert-pli: Modeling paragraph-level interactions for legal case retrieval
Yunqiu Shao, Jiaxin Mao, Yiqun Liu, Weizhi Ma, Ken Satoh, Min Zhang, and Shaoping Ma. 2020 · 2020
Earlier work this paper cites.
Ledgar: A large-scale multi-label corpus for text classification of legal provisions in contracts
Don Tuggener, Pius Von Däniken, Thomas Peetz, and Mark Cieliebak. 2020 · 2020
Cited alongside, same era.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020 · 2020
Cited alongside, same era.
Incorporating domain knowledge for extractive summarization of legal case documents
Paheli Bhattacharya, Soham Poddar, Koustav Rudra, Kripabandhu Ghosh, and Saptarshi Ghosh. 2021 · 2021
Cited alongside, same era.
Indonlg: Benchmark and resources for evaluating indonesian natural language generation
Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Khodra, et al. 2021 · 2021
Cited alongside, same era.
Multieurlex-a multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer
Ilias Chalkidis, Manos Fergadiotis, and Ion Androutsopoulos. 2021 · 2021
Indicnlg benchmark: Multilingual datasets for diverse nlg tasks in indic languages
Aman Kumar, Himani Shrotriya, Prachi Sahu, Amogh Mishra, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, Mitesh M Khapra, and Pratyush Kumar. 2022 · 2022
Later among the works it cites.
Italian-legal-bert: A pre-trained transformer language model for italian law
Daniele Licari and Giovanni Comandè. 2022 · 2022
Later among the works it cites.
A statutory article retrieval dataset in french
Antoine Louis and Gerasimos Spanakis. 2022 · 2022
Later among the works it cites.
M3: Multi-level dataset for multi-document summarisation of medical studies
Julia Otmakhova, Karin Verspoor, Timothy Baldwin, Antonio Jimeno Yepes, and Jey Han Lau. 2022 · 2022
Later among the works it cites.
Abstractive summarization of dutch court verdicts using sequence-to-sequence models
Marijn Schraagen, Floris Bex, Nick Van De Luijtgaarden, and Daniël Prijs. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bertbr: a pretrained language model for law texts
Victor Hugo Ciurlino. 2021 · 2021
Cited alongside, same era.
Measuring law over time: A network analytical framework with an application to statutes and regulations in the united states and germany
Corinna Coupette, Janis Beckedorf, Dirk Hartung, Michael Bommarito, and Daniel Martin Katz. 2021 · 2021
Cited alongside, same era.
Juribert: A masked-language model adaptation for french legal text
Stella Douka, Hadi Abdine, Michalis Vazirgiannis, Rajaa El Hamdani, and David Restrepo Amariles. 2021 · 2021
Cited alongside, same era.
The gem benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna Clinciu, Dipanjan Das, Kaustubh Dhole, et al. 2021 · 2021
Cited alongside, same era.
Spanish legalese language model and corpora
Asier Gutiérrez-Fandiño, Jordi Armengol-Estapé, Aitor Gonzalez-Agirre, and Marta Villegas. 2021 · 2021
Cited alongside, same era.
Cuad: An expert-annotated nlp dataset for legal contract review
Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball. 2021 · 2021
Cited alongside, same era.
Efficient attentions for long document summarization
Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang. 2021 · 2021
Cited alongside, same era.
Scrolls: Standardized comparison over long language sequences
Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, et al. 2022 · 2022
Later among the works it cites.
Multi-lexsum: Real-world summaries of civil rights lawsuits at multiple granularities
Zejiang Shen, Kyle Lo, Lauren Yu, Nathan Dahlberg, Margo Schlanger, and Doug Downey. 2022 · 2022
Later among the works it cites.
Legal case document summarization: Extractive and abstractive methods and their evaluation
Abhay Shukla, Paheli Bhattacharya, Soham Poddar, Rajdeep Mukherjee, Kripabandhu Ghosh, Pawan Goyal, and Saptarshi Ghosh. 2022 · 2022
Later among the works it cites.
Lamberta: Law article mining based on bert architecture for the italian civil code
Andrea Tagarelli and Andrea Simeri. 2022 · 2022
Later among the works it cites.
Primera: Pyramid-based masked sentence pre-training for multi-document summarization
Wen Xiao, Iz Beltagy, Giuseppe Carenini, and Arman Cohan. 2022 · 2022
Later among the works it cites.
Unlimiformer: Long-range transformers with unlimited length input
Amanda Bertsch, Uri Alon, Graham Neubig, and Matthew R Gormley. 2023 · 2023
Later among the works it cites.
Banglanlg and banglat5: Benchmarks and resources for evaluating low-resource natural language generation in bangla
Abhik Bhattacharjee, Tahmid Hasan, Wasi Ahmad, and Rifat Shahriyar. 2023 · 2023
Later among the works it cites.
LeXFiles and LegalLAMA: Facilitating English multinational legal language model development
Ilias Chalkidis, Nicolas Garneau, Catalina Goanta, Daniel Katz, and Anders Søgaard. 2023 · 2023
Later among the works it cites.
Booookscore: A systematic exploration of book-length summarization in the era of llms
Yapei Chang, Kyle Lo, Tanya Goyal, and Mohit Iyyer. 2023 · 2023
Later among the works it cites.
Equals: A real-world dataset for legal question answering via reading chinese laws
Andong Chen, Feng Yao, Xinyan Zhao, Yating Zhang, Changlong Sun, Yun Liu, and Weixing Shen. 2023 · 2023
Later among the works it cites.
Towards argument-aware abstractive summarization of long legal opinions with summary reranking
Mohamed Elaraby, Yang Zhong, and Diane Litman. 2023 · 2023
Later among the works it cites.
Dolphin: A challenging and diverse benchmark for arabic nlg
Abdelrahim Elmadany, Ahmed El-Shangiti, Muhammad Abdul-Mageed, et al. 2023 · 2023
Later among the works it cites.
Lawbench: Benchmarking legal knowledge of large language models
Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, Zhuo Han, Songyang Zhang, Kai Chen, Zongwen Shen, and Jidong Ge. 2023 · 2023
Later among the works it cites.
Natural language processing for legal document review: categorising deontic modalities in contracts
S Georgette Graham, Hamidreza Soltani, and Olufemi Isiaq. 2023 · 2023
Later among the works it cites.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Neel Guha, Julian Nyarko, Daniel E Ho, Christopher Re, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, et al. 2023 · 2023
Later among the works it cites.
Medeval: A multi-level, multi-task, and multi-domain medical benchmark for language model evaluation
Zexue He, Yu Wang, An Yan, Yao Liu, Eric Chang, Amilcare Gentili, Julian McAuley, and Chun-nan Hsu. 2023 · 2023
Later among the works it cites.
Efficient long-text understanding with short-text models
Maor Ivgi, Uri Shaham, and Jonathan Berant. 2023 · 2023
Later among the works it cites.
Natural language processing in the legal domain
Daniel Martin Katz, Dirk Hartung, Lauritz Gerlach, Abhik Jana, and Michael James Bommarito. 2023 · 2023
Later among the works it cites.
An exploration of encoder-decoder approaches to multi-label classification for legal and biomedical text
Yova Kementchedjhieva and Ilias Chalkidis. 2023 · 2023
Later among the works it cites.
Interpretable long-form legal question answering with retrieval-augmented large language models
Antoine Louis, Gijs van Dijck, and Gerasimos Spanakis. 2023 · 2023
Later among the works it cites.
Pre-trained language models for the legal domain: a case study on indian law
Shounak Paul, Arpan Mandal, Pawan Goyal, and Saptarshi Ghosh. 2023 · 2023
Later among the works it cites.
What to read in a contract? party-specific summarization of legal obligations, entitlements, and prohibitions
Abhilasha Sancheti, Aparna Garimella, Balaji Srinivasan, and Rachel Rudinger. 2023 · 2023
Later among the works it cites.
Maud: An expert-annotated legal nlp dataset for merger agreement understanding
Steven H Wang, Antoine Scardigli, Leonard Tang, Wei Chen, Dimitry Levkin, Anya Chen, Spencer Ball, Thomas Woodside, Oliver Zhang, and Dan Hendrycks. 2023 · 2023
Later among the works it cites.
Argumentative segmentation enhancement for legal summarization
Huihui Xu and Kevin Ashley. 2023 · 2023
Later among the works it cites.
Lexabsumm: Aspect-based summarization of legal decisions
Santosh Tyss, Mahmoud Aly, and Matthias Grabmair. 2024 · 2024
Closest in time.