Fetching the paper…
Reading the bibliography…
In the rapidly evolving field of large language models (LLMs), data augmentation (DA) has emerged as a pivotal technique for enhancing model performance by diversifying training examples without the need for additional data collection.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A Raffel. 2022 · 1965
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Wearable devices in healthcare: Privacy and information security issues
Liezel Cilliers. 2019 · 2019
Earlier work this paper cites.
Efficient intent detection with dual sentence encoders
Iñigo Casanueva, Tadas Temčinas, Daniela Gerz, Matthew Henderson, and Ivan Vulić. 2020 · 2020
Earlier work this paper cites.
Recommendation system based on deep learning methods: a systematic review and new directions
Aminu Da’u and Naomie Salim. 2020 · 2020
Earlier work this paper cites.
A survey on knowledge graph-based recommender systems
Qingyu Guo, Fuzhen Zhuang, Chuan Qin, Hengshu Zhu, Xing Xie, Hui Xiong, and Qing He. 2020 · 2020
Earlier work this paper cites.
Tell me how to ask again: Question data augmentation with controllable rewriting in continuous space
Dayiheng Liu, Yeyun Gong, Jie Fu, Yu Yan, Jiusheng Chen, Jiancheng Lv, Nan Duan, and M. Zhou. 2020 · 2020
Earlier work this paper cites.
Improving conversational recommender systems via knowledge graph based semantic fusion
Kun Zhou, Wayne Xin Zhao, Shuqing Bian, Yuanhang Zhou, Ji-Rong Wen, and Jingsong Yu. 2020 · 2020
Earlier work this paper cites.
Medically aware gpt-3 as a data generator for medical dialogue summarization
Bharath Chintagunta, Namit Katariya, Xavier Amatriain, and Anitha Kannan. 2021 · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021 · 2021
Earlier work this paper cites.
A survey of data augmentation approaches for nlp
Steven Y. Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard H. Hovy. 2021 · 2021
Earlier work this paper cites.
Re-ranking with constraints on diversified exposures for homepage recommender system
Qi Hao, Tianze Luo, and Guangda Huzhang. 2021 · 2021
Earlier work this paper cites.
A survey on recent approaches for natural language processing in low-resource scenarios
Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, and Dietrich Klakow. 2021 · 2021
Earlier work this paper cites.
Want to reduce labeling cost? gpt-3 can help
Shuohang Wang, Yang Liu, Yichong Xu, Chenguang Zhu, and Michael Zeng. 2021 · 2021
Earlier work this paper cites.
GPT3Mix: Leveraging large-scale language models for text augmentation
Kang Min Yoo, Dongju Park, Jaewook Kang, Sang-Woo Lee, and Woomyoung Park. 2021 · 2021
Earlier work this paper cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Earlier work this paper cites.
Inductive biases for low data vqa: a data augmentation approach
Narjes Askarian, Ehsan Abbasnejad, Ingrid Zukerman, Wray Buntine, and Gholamreza Haffari. 2022 · 2022
Earlier work this paper cites.
Ew-tune: A framework for privately fine-tuning large language models with differential privacy
Rouzbeh Behnia, Mohammadreza Reza Ebrahimi, Jason Pacheco, and Balaji Padmanabhan. 2022 · 2022
Earlier work this paper cites.
Inpars: Data augmentation for information retrieval using large language models
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022 · 2022
Earlier work this paper cites.
M6-rec: Generative pretrained language models are open-ended recommender systems
Zeyu Cui, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022 · 2022
Earlier work this paper cites.
CORE: A retrieve-then-edit framework for counterfactual data generation
Tanay Dixit, Bhargavi Paranjape, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
Datscore: Evaluating translation with data augmented translations
Moussa Kamal Eddine, Guokan Shang, and Michalis Vazirgiannis. 2022 · 2022
Earlier work this paper cites.
Large language models are reasoning teachers
Namgyu Ho, Laura Schmid, and Se-Young Yun. 2022 · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and L. Sifre. 2022 · 2022
Earlier work this paper cites.
Teacher-student architecture for knowledge learning: A survey
Chengming Hu, Xuan Li, Dan Liu, Xi Chen, Ju Wang, and Xue Liu. 2022 · 2022
Earlier work this paper cites.
Score-based generative modeling of graphs via the system of stochastic differential equations
Jaehyeong Jo, Seul Lee, and Sung Ju Hwang. 2022 · 2022
Earlier work this paper cites.
Tackling covid-19 conspiracies on twitter using bert ensembles, gpt-3 augmentation, and graph nns
Damir Korenčić, Ivan Grubišić, Gretel Liz De La Peña Sarracén, Alejandro Hector Toselli, Berta Chulvi, and Paolo Rosso. 2022 · 2022
Earlier work this paper cites.
Controllable dialogue simulation with in-context learning
Zekun Li, Wenhu Chen, Shiyang Li, Hong Wang, Jing Qian, and Xifeng Yan. 2022a · 2022
Earlier work this paper cites.
Data contamination: From memorization to exploitation
Inbal Magar and Roy Schwartz. 2022 · 2022
Earlier work this paper cites.
European clinical case corpus
Bernardo Magnini, Begoña Altuna, Alberto Lavelli, Anne-Lyse Minard, Manuela Speranza, and Roberto Zanoli. 2022 · 2022
Earlier work this paper cites.
Will we run out of data? an analysis of the limits of scaling datasets in machine learning
Pablo Villalobos, Jaime Sevilla, Lennart Heim, Tamay Besiroglu, Marius Hobbhahn, and Anson Ho. 2022 · 2022
Earlier work this paper cites.
A unified dialogue user simulator for few-shot data augmentation
Dazhen Wan, Zheng Zhang, Qi Zhu, Lizi Liao, and Minlie Huang. 2022a · 2022
Earlier work this paper cites.
A unified dialogue user simulator for few-shot data augmentation
Dazhen Wan, Zheng Zhang, Qi Zhu, Lizi Liao, and Minlie Huang. 2022b · 2022
Earlier work this paper cites.
Meysam Alizadeh, Maël Kubli, Zeynab Samei, Shirin Dehghani, Juan Diego Bermeo, Maria Korobeynikova, and Fabrizio Gilardi. 2023 · 2023
Earlier work this paper cites.
Out of one, many: Using language models to simulate human samples
Lisa Argyle, Ethan Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, and David Wingate. 2023 · 2023
Earlier work this paper cites.
Large language models as annotators: Enhancing generalization of nlp models at minimal cost
Parikshit Bansal and Amit Sharma. 2023 · 2023
Earlier work this paper cites.
A systematic evaluation of large language models on out-of-distribution logical reasoning tasks
Qiming Bao, Gael Gendron, Alex Yuxuan Peng, Wanjun Zhong, Neset Tan, Yang Chen, Michael Witbrock, and Jiamou Liu. 2023 · 2023
Earlier work this paper cites.
Leveraging chatgpt as text annotation tool for sentiment analysis
Mohammad Belal, James She, and Simon Wong. 2023 · 2023
Earlier work this paper cites.
Thales Bertaglia, Stefan Huber, Catalina Goanta, Gerasimos Spanakis, and Adriana Iamnitchi. 2023 · 2023
Earlier work this paper cites.
Fine-tuning a llm using reinforcement learning from human feedback for a therapy chatbot application
Desirée Bill and Theodor Eriksson. 2023 · 2023
Earlier work this paper cites.
Synthesizing mixed-type electronic health records using diffusion models
Taha Ceritli, Ghadeer O Ghosheh, Vinod Kumar Chauhan, Tingting Zhu, Andrew P Creagh, and David A Clifton. 2023 · 2023
Earlier work this paper cites.
Chain-of-thought prompt distillation for multimodal named entity and multimodal relation extraction
Feng Chen and Yujian Feng. 2023 · 2023
Cited alongside, same era.
Large language models for user interest journeys
Konstantina Christakopoulou, Alberto Lalama, Cj Adams, Iris Qu, Yifat Amir, Samer Chucri, Pierce Vollucci, Fabio Soldo, Dina Bseiso, Sarah Scodel, et al. 2023 · 2023
Cited alongside, same era.
Data-centric financial large language models
Zhixuan Chu, Huaiyu Guo, Xinyuan Zhou, Yijia Wang, Fei Yu, Hong Chen, Wanqing Xu, Xin Lu, Qing Cui, Longfei Li, Jun Zhou, and Sheng Li. 2023 · 2023
Cited alongside, same era.
Chataug: Leveraging chatgpt for text data augmentation
Haixing Dai, Zhengliang Liu, Wenxiong Liao, Xiaoke Huang, Zihao Wu, Lin Zhao, Wei Liu, Ninghao Liu, Sheng Li, Dajiang Zhu, et al. 2023 · 2023
Cited alongside, same era.
Charles O’Neill, Yuan-Sen Ting, Ioana Ciuca, Roberta Raileanu, Jack Miller, and Thang Bui. 2023 · 2023
Later among the works it cites.
Keivalya Pandya and Mehfuza Holia. 2023 · 2023
Later among the works it cites.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023 · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hexuan Deng, Xin Zhang, Meishan Zhang, Xuebo Liu, and Min Zhang. 2023 · 2023
Cited alongside, same era.
Diversify your vision datasets with automatic diffusion-based augmentation
Lisa Dunlap, Alyssa Umino, Han Zhang, Jiezhi Yang, Joseph E Gonzalez, and Trevor Darrell. 2023 · 2023
Cited alongside, same era.
Recommender systems in the era of large language models (llms)
Wenqi Fan, Zihuai Zhao, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Jiliang Tang, and Qing Li. 2023 · 2023
Cited alongside, same era.
Chatgpt as data augmentation for compositional generalization: A case study in open intent detection
Yihao Fang, Xianzhi Li, Stephen W Thomas, and Xiaodan Zhu. 2023 · 2023
Cited alongside, same era.
A survey of graph neural networks for recommender systems: Challenges, methods, and directions
Chen Gao, Yu Zheng, Nian Li, Yinfeng Li, Yingrong Qin, Jinghua Piao, Yuhan Quan, Jianxin Chang, Depeng Jin, Xiangnan He, et al. 2023 · 2023
Cited alongside, same era.
Dale: Generative data augmentation for low-resource legal nlp
Sreyan Ghosh, Chandra Kiran Evuru, Sonal Kumar, S Ramaneswaran, S Sakshi, Utkarsh Tyagi, and Dinesh Manocha. 2023 · 2023
Cited alongside, same era.
Chatgpt outperforms crowd workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Cited alongside, same era.
Large language models respond to influence like humans
Lewis Griffin, Bennett Kleinberg, Maximilian Mozes, Kimberly Mai, Maria Do Mar Vau, Matthew Caldwell, and Augustine Mavor-Parker. 2023 · 2023
Cited alongside, same era.
Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark
Oscar Sainz, Jon Ander Campos, Iker García-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre. 2023 · 2023
Later among the works it cites.
Can llms augment low-resource reading comprehension datasets? opportunities and challenges
Vinay Samuel, Houda Aynaou, Arijit Ghosh Chowdhury, Karthik Venkat Ramanan, and Aman Chadha. 2023 · 2023
Later among the works it cites.
Shouvon Sarker, Lijun Qian, and Xishuang Dong. 2023 · 2023
Later among the works it cites.
Viktor Schlegel, Hao Li, Yuping Wu, Anand Subramanian, Thanh-Tung Nguyen, Abhinav Ramesh Kashyap, Daniel Beck, Xiaojun Zeng, Riza Theresa Batista-Navarro, Stefan Winkler, et al. 2023 · 2023
Later among the works it cites.
Automatic prompt augmentation and selection with chain-of-thought from labeled data
KaShun Shum, Shizhe Diao, and Tong Zhang. 2023 · 2023
Later among the works it cites.
The curse of recursion: Training on generated data makes models forget
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson. 2023 · 2023
Later among the works it cites.
From fake to hyperpartisan news detection using domain adaptation
Răzvan-Alexandru Smădu, Sebastian-Vasile Echim, Dumitru-Clementin Cercel, Iuliana Marin, and Florin Pop. 2023 · 2023
Later among the works it cites.
Large language models for enhancing customer lifecycle management
Vishvesh Soni. 2023 · 2023
Later among the works it cites.
Improving robustness in knowledge distillation using domain-targeted data augmentation
Joe Stacey and Marek Rei. 2023 · 2023
Later among the works it cites.
Just-in-time security patch detection–llm at the rescue for data augmentation
Xunzhu Tang, Zhenghan Chen, Kisub Kim, Haoye Tian, Saad Ezzini, and Jacques Klein. 2023 · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Later among the works it cites.
Petter Törnberg. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Later among the works it cites.
Petter Törnberg. 2023 · 2023
Later among the works it cites.
Llm-powered data augmentation for enhanced crosslingual performance
Chenxi Whitehouse, Monojit Choudhury, and Alham Fikri Aji. 2023 · 2023
Later among the works it cites.
Multimodal data augmentation for image captioning using diffusion models
Changrong Xiao, Sean Xin Xu, and Kunpeng Zhang. 2023 · 2023
Later among the works it cites.
Empowering llm-based machine translation with cultural awareness
Binwei Yao, Ming Jiang, Diyi Yang, and Junjie Hu. 2023 · 2023
Later among the works it cites.
From playing the story to gaming the system: Repeat experiences of a large language model-based interactive story
Qing Ru Yong and Alex Mitchell. 2023 · 2023
Later among the works it cites.
Data-centric artificial intelligence: A survey
Daochen Zha, Zaid Pervaiz Bhat, Kwei-Herng Lai, Fan Yang, Zhimeng Jiang, Shaochen Zhong, and Xia Hu. 2023 · 2023
Later among the works it cites.
Ask an expert: Leveraging language models to improve strategic reasoning in goal-oriented dialogue models
Qiang Zhang, Jason Naradowsky, and Yusuke Miyao. 2023e · 2023
Later among the works it cites.
AugESC: Dialogue augmentation with large language models for emotional support conversation
Chujie Zheng, Sahand Sabour, Jiaxin Wen, Zheng Zhang, and Minlie Huang. 2023a · 2023
Later among the works it cites.
Augesc: Dialogue augmentation with large language models for emotional support conversation
Chujie Zheng, Sahand Sabour, Jiaxin Wen, Zheng Zhang, and Minlie Huang. 2023b · 2023
Later among the works it cites.
Somnath Banerjee, Amruit Sahoo, Sayan Layek, Avik Dutta, Rima Hazra, and Animesh Mukherjee. 2024 · 2024
Closest in time.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Homes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Wing Yin Ng, Ricky Wang, and Aditya Ramesh. 2024 · 2024
Closest in time.
Drlc: Reinforcement learning with dense rewards from llm critic
Meng Cao, Lei Shu, Lei Yu, Yun Zhu, Nevan Wichers, Yinxiao Liu, and Lei Meng. 2024 · 2024
Closest in time.
Diversify your vision datasets with automatic diffusion-based augmentation
Lisa Dunlap, Alyssa Umino, Han Zhang, Jiezhi Yang, Joseph E Gonzalez, and Trevor Darrell. 2024 · 2024
Closest in time.
Large language models to identify social determinants of health in electronic health records
Marco Guevara, Shan Chen, Spencer Thomas, Tafadzwa L Chaunzwa, Idalid Franco, Benjamin H Kann, Shalini Moningi, Jack M Qian, Madeleine Goldstein, Susan Harper, et al. 2024 · 2024
Closest in time.
Mina J Kian, Mingyu Zong, Katrin Fischer, Abhyuday Singh, Anna-Maria Velentza, Pau Sang, Shriya Upadhyay, Anika Gupta, Misha A Faruki, Wallace Browning, et al. 2024 · 2024
Closest in time.
Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities
Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, and Bryan Catanzaro. 2024 · 2024
Closest in time.
Graph principal flow network for conditional graph generation
Zhanfeng Mo, Tianze Luo, and Sinno Jialin Pan. 2024 · 2024
Closest in time.
Unifying large language models and knowledge graphs: A roadmap
Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu. 2024 · 2024
Closest in time.
Llms in e-commerce: a comparative analysis of gpt and llama models in product review evaluation
Konstantinos I Roumeliotis, Nikolaos D Tselikas, and Dimitrios K Nasiopoulos. 2024 · 2024
Closest in time.
Yiping Song, Juhua Zhang, Zhiliang Tian, Yuxin Yang, Minlie Huang, and Dongsheng Li. 2024 · 2024
Closest in time.
Weaver: Foundation models for creative writing
Tiannan Wang, Jiamin Chen, Qingrui Jia, Shuai Wang, Ruoyu Fang, Huilin Wang, Zhaowei Gao, Chunzhao Xie, Chuou Xu, Jihong Dai, et al. 2024 · 2024
Closest in time.
Large language models can learn temporal reasoning
Siheng Xiong, Ali Payani, Ramana Kompella, and Faramarz Fekri. 2024 · 2024
Closest in time.
Gpt4tools: Teaching large language model to use tools via self-instruction
Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, and Ying Shan. 2024 · 2024
Closest in time.
Large language models for social networks: Applications, challenges, and solutions
Jingying Zeng, Richard Huang, Waleed Malik, Langxuan Yin, Bojan Babic, Danny Shacham, Xiao Yan, Jaewon Yang, and Qi He. 2024 · 2024
Closest in time.
What makes good examples for visual in-context learning?
Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. 2024 · 2024
Closest in time.