Fetching the paper…
Reading the bibliography…
We present DALE, a novel and effective generative Data Augmentation framework for low-resource LEgal NLP.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Eda: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou. 2019 · 1901
Earlier work this paper cites.
Augmenting data with mixup for sentence classification: An empirical study
Hongyu Guo, Yongyi Mao, and Richong Zhang. 2019 · 1905
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
ContractNLI: A dataset for document-level natural language inference for contracts
Yuta Koreeda and Christopher Manning. 2021 · 1919
Earlier work this paper cites.
Transmission of information: A statistical theory of communications
Robert M Fano. 1961 · 1961
Earlier work this paper cites.
Perplexity—a measure of the difficulty of speech recognition tasks
Fred Jelinek, Robert L Mercer, Lalit R Bahl, and James K Baker. 1977 · 1977
Earlier work this paper cites.
Excursions into the nature of legal language
Mary Jane Morrison. 1989 · 1989
Earlier work this paper cites.
Wordnet: a lexical database for english
George A Miller. 1995 · 1995
Earlier work this paper cites.
The pagerank citation ranking: Bringing order to the web
Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd. 1999 · 1999
Earlier work this paper cites.
Discovering word senses from text
Patrick Pantel and Dekang Lin. 2002 · 2002
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
How does nlp benefit legal system: A summary of legal artificial intelligence
Haoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang, Zhiyuan Liu, and Maosong Sun. 2020a · 2004
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn et al. 2005 · 2005
Earlier work this paper cites.
What are the productive units of natural language grammar? a dop approach to the automatic identification of constructions
Willem Zuidema. 2006 · 2006
Earlier work this paper cites.
Christopher williams, tradition and change in legal english: Verbal constructions in prescriptive texts
Ann Sinsheimer. 2007 · 2007
Earlier work this paper cites.
Tradition and Change in Legal English: Verbal Constructions in Prescriptive Texts , volume 20
Christopher Williams. 2007 · 2007
Earlier work this paper cites.
Ssmba: Self-supervised manifold based data augmentation for improving out-of-domain robustness
Nathan Ng, Kyunghyun Cho, and Marzyeh Ghassemi. 2020a · 2009
Earlier work this paper cites.
Legal-bert: The muppets straight out of law school
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos. 2020 · 2010
Earlier work this paper cites.
Mixup-transformer: dynamic data augmentation for nlp tasks
Lichao Sun, Congying Xia, Wenpeng Yin, Tingting Liang, Philip S Yu, and Lifang He. 2020 · 2010
Earlier work this paper cites.
A broad evaluation of techniques for automatic acquisition of multiword expressions
Carlos Ramisch, Vitor De Araujo, and Aline Villavicencio. 2012 · 2012
Earlier work this paper cites.
Supreme court database, version 2013 release 01
Harold J Spaeth, Lee Epstein, Andrew D Martin, Jeffrey A Segal, Theodore J Ruger, and Sara C Benesh. 2013 · 2013
Earlier work this paper cites.
Taking the best from the crowd:learning question passage classification from noisy data
Azad Abad and Alessandro Moschitti. 2016 · 2016
Earlier work this paper cites.
Aggregating and predicting sequence labels from crowd annotations
An Thanh Nguyen, Byron Wallace, Junyi Jessy Li, Ani Nenkova, and Matthew Lease. 2017 · 2017
Earlier work this paper cites.
CLAUDETTE: an automated detector of potentially unfair clauses in online terms of service
Marco Lippi, Przemyslaw Palka, Giuseppe Contissa, Francesca Lagioia, Hans-Wolfgang Micklitz, Giovanni Sartor, and Paolo Torroni. 2018 · 2018
Earlier work this paper cites.
Legal document retrieval using document vector embeddings and deep learning
Keet Sugathadasa, Buddhi Ayesha, Nisansa de Silva, Amal Shehan Perera, Vindula Jayawardana, Dimuthu Lakmal, and Madhavi Perera. 2019 · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Earlier work this paper cites.
Qanet: Combining local convolution with global self-attention for reading comprehension
Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, and Quoc V Le. 2018 · 2018
Earlier work this paper cites.
FLAIR: An easy-to-use framework for state-of-the-art NLP
Alan Akbik, Tanja Bergmann, Duncan Blythe, Kashif Rasul, Stefan Schweter, and Roland Vollgraf. 2019 · 2019
Earlier work this paper cites.
Neural legal judgment prediction in English
Ilias Chalkidis, Ion Androutsopoulos, and Nikolaos Aletras. 2019 · 2019
Earlier work this paper cites.
Keep calm and switch on! preserving sentiment and fluency in semantic text exchange
Steven Y. Feng, Aaron W. Li, and Jesse Hoey. 2019 · 2019
Earlier work this paper cites.
Fine-grained named entity recognition in legal documents
Elena Leitner, Georg Rehm, and Julian Moreno-Schneider. 2019 · 2019
Earlier work this paper cites.
Claudette: an automated detector of potentially unfair clauses in online terms of service
Marco Lippi, Przemysław Pałka, Giuseppe Contissa, Francesca Lagioia, Hans-Wolfgang Micklitz, Giovanni Sartor, and Paolo Torroni. 2019 · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019 · 2019
Cited alongside, same era.
Predicting annotation difficulty to improve task routing and model performance for biomedical information extraction
Yinfei Yang, Oshin Agarwal, Chris Tar, Byron C. Wallace, and Ani Nenkova. 2019 · 2019
Cited alongside, same era.
Tax Law NLP Resources
Andrew Blair-Stanek, Nils Holzenberger, and Benjamin Van Durme. 2020 · 2020
Cited alongside, same era.
Contract discovery: Dataset and a few-shot semantic retrieval challenge with competitive baselines
Łukasz Borchmann, Dawid Wisniewski, Andrzej Gretkowski, Izabela Kosmala, Dawid Jurkiewicz, Łukasz Szałkiewicz, Gabriela Pałka, Karol Kaczmarek, Agnieszka Kaliska, and Filip Graliński. 2020 · 2020
Style transfer as data augmentation: A case study on named entity recognition
Shuguang Chen, Leonardo Neves, and Thamar Solorio. 2022 · 2022
Later among the works it cites.
LongT5: Efficient text-to-text transformer for long sequences
Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang. 2022b · 2022
Later among the works it cites.
Pile of law: Learning responsible data filtering from the law and a 256gb open-source legal dataset
Peter Henderson, Mark Krass, Lucia Zheng, Neel Guha, Christopher D Manning, Dan Jurafsky, and Daniel Ho. 2022 · 2022
Later among the works it cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022 · 2022
Later among the works it cites.
Named entity recognition in Indian court judgments
Prathamesh Kalamkar, Astha Agarwal, Aman Tiwari, Smita Gupta, Saurabh Karn, and Vivek Raghavan. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An analysis of simple data augmentation for named entity recognition
Xiang Dai and Heike Adel. 2020 · 2020
Cited alongside, same era.
DAGA: Data augmentation with a generation approach for low-resource tagging tasks
Bosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai, Thien Hai Nguyen, Shafiq Joty, Luo Si, and Chunyan Miao. 2020 · 2020
Cited alongside, same era.
Nonlinear mixup: Out-of-manifold data augmentation for text classification
Hongyu Guo. 2020 · 2020
Cited alongside, same era.
UMLS-based data augmentation for natural language processing of clinical research literature
Tian Kang, Adler Perotte, Youlan Tang, Casey Ta, and Chunhua Weng. 2020 · 2020
Cited alongside, same era.
Data augmentation using pre-trained transformer models
Varun Kumar, Ashutosh Choudhary, and Eunah Cho. 2020 · 2020
Cited alongside, same era.
How effective is task-agnostic data augmentation for pretrained transformers?
Shayne Longpre, Yu Wang, and Chris DuBois. 2020 · 2020
Cited alongside, same era.
SSMBA: Self-supervised manifold based data augmentation for improving out-of-domain robustness
Nathan Ng, Kyunghyun Cho, and Marzyeh Ghassemi. 2020b · 2020
Cited alongside, same era.
Alp: Data augmentation using lexicalized pcfgs for few-shot text classification
Hazel H Kim, Daecheol Woo, Seong Joon Oh, Jeong-Won Cha, and Yo-Sub Han. 2022 · 2022
Later among the works it cites.
Semantic segmentation of legal documents via rhetorical roles
Vijit Malik, Rishabh Sanjay, Shouvik Kumar Guha, Angshuman Hazarika, Shubham Nigam, Arnab Bhattacharya, and Ashutosh Modi. 2022 · 2022
Later among the works it cites.
Processing long legal documents with pre-trained transformers: Modding LegalBERT and longformer
Dimitris Mamakas, Petros Tsotsi, Ion Androutsopoulos, and Ilias Chalkidis. 2022 · 2022
Later among the works it cites.
Budgetlongformer: Can we cheaply pretrain a sota legal language model from scratch?
Joel Niklaus and Daniele Giofré. 2022 · 2022
Later among the works it cites.
Combining wordnet and word embeddings in data augmentation for legal texts
Sezen Perçin, Andrea Galassi, Francesca Lagioia, Federico Ruggeri, Piera Santin, Giovanni Sartor, and Paolo Torroni. 2022 · 2022
Later among the works it cites.
InforMask: Unsupervised informative masking for language model pretraining
Nafis Sadeq, Canwen Xu, and Julian McAuley. 2022 · 2022
Later among the works it cites.
Legal case document summarization: Extractive and abstractive methods and their evaluation
Abhay Shukla, Paheli Bhattacharya, Soham Poddar, Rajdeep Mukherjee, Kripabandhu Ghosh, Pawan Goyal, and Saptarshi Ghosh. 2022 · 2022
Later among the works it cites.
PromDA: Prompt-based data augmentation for low-resource NLU tasks
Yufei Wang, Can Xu, Qingfeng Sun, Huang Hu, Chongyang Tao, Xiubo Geng, and Daxin Jiang. 2022 · 2022
Later among the works it cites.
Caselaw access project
2018 · 2023
Closest in time.
Chatgpt may pass the bar exam soon, but has a long way to go for the lexglue benchmark
Ilias Chalkidis. 2023 · 2023
Closest in time.
LeXFiles and LegalLAMA: Facilitating English Multinational Legal Language Model Development
Ilias Chalkidis*, Nicolas Garneau*, Catalina Goanta, Daniel Martin Katz, and Anders Søgaard. 2023 · 2023
Closest in time.
An empirical survey of data augmentation for limited data learning in nlp
Jiaao Chen, Derek Tam, Colin Raffel, Mohit Bansal, and Diyi Yang. 2023 · 2023
Closest in time.
Chataug: Leveraging chatgpt for text data augmentation
Haixing Dai, Zheng Liu, Wenxiong Liao, Xiaoke Huang, Zihao Wu, Lin Zhao, Wei Liu, Ninghao Liu, Sheng Li, Dajiang Zhu, Hongmin Cai, Quanzheng Li, Dinggang Shen, Tianming Liu, and Xiang Li. 2023 · 2023
Closest in time.
Decoder-only or encoder-decoder? interpreting language model as a regularized encoder-decoder
Zihao Fu, Wai Lam, Qian Yu, Anthony Man-Cho So, Shengding Hu, Zhiyuan Liu, and Nigel Collier. 2023 · 2023
Closest in time.
How much data are augmentations worth? an investigation into scaling laws, invariance, and implicit regularization
Jonas Geiping, Micah Goldblum, Gowthami Somepalli, Ravid Shwartz-Ziv, Tom Goldstein, and Andrew Gordon Wilson. 2023 · 2023
Closest in time.
Bioaug: Conditional generation based data augmentation for low-resource biomedical ner
Sreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, and Dinesh Manocha. 2023 · 2023
Closest in time.
The concept of legal language: What makes legal language ‘legal ‘?
Ondřej Glogar. 2023 · 2023
Closest in time.
International Legal English: A Practical Introduction for Students and Professionals
Rupert Haigh. 2023 · 2023
Closest in time.
Huggingfaceh4/open_llm_leaderboard
HuggingFace. 2023 · 2023
Closest in time.
Efficient long-text understanding with short-text models
Maor Ivgi, Uri Shaham, and Jonathan Berant. 2023 · 2023
Closest in time.
Natural language processing in the legal domain
Daniel Martin Katz, Dirk Hartung, Lauritz Gerlach, Abhik Jana, and Michael J Bommarito II. 2023 · 2023
Closest in time.
Automatic rhetorical roles classification for legal documents using legal-transformeroverbert
Gabriele Marino, Daniele Licari, Praveen Bushipaka, Giovanni Comandé, and Tommaso Cucinotta. 2023 · 2023
Closest in time.
Exploiting language characteristics for legal domain-specific language model pretraining
Inderjeet Nair and Natwar Modani. 2023 · 2023
Closest in time.
Lextreme: A multi-lingual and multi-task benchmark for the legal domain
Joel Niklaus, Veton Matoshi, Pooja Rani, Andrea Galassi, Matthias Stürmer, and Ilias Chalkidis. 2023 · 2023
Closest in time.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023 · 2023
Closest in time.
Ul2: Unifying language learning paradigms
Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, et al. 2023 · 2023
Closest in time.
Maud: An expert-annotated legal nlp dataset for merger agreement understanding
Steven H. Wang, Antoine Scardigli, Leonard Tang, Wei Chen, Dimitry Levkin, Anya Chen, Spencer Ball, Thomas Woodside, Oliver Zhang, and Dan Hendrycks. 2023 · 2023
Closest in time.