Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) excel at generating synthetic data, but ensuring its quality and diversity remains challenging.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Semeval-2010 task 8: Multi-way classification of semantic relations between pairs of nominals
Iris Hendrickx, Su Nam Kim, Zornitsa Kozareva, Preslav Nakov, Diarmuid Ó Séaghdha, Sebastian Padó, Marco Pennacchiotti, Lorenza Romano, and Stan Szpakowicz. 2019 · 1911
Earlier work this paper cites.
The need for biases in learning generalizations
Tom M Mitchell. 1980 · 1980
Earlier work this paper cites.
A linear programming formulation for global inference in natural language tasks
Dan Roth and Wen-tau Yih. 2004 · 2004
Earlier work this paper cites.
ChemProt: a disease chemical biology database
Olivier Taboureau, Sonny Kim Nielsen, Karine Audouze, Nils Weinhold, Daniel Edsgärd, Francisco S. Roque, Irene Kouskoumvekaki, Alina Bora, Ramona Curpan, Thomas Skøt Jensen, Søren Brunak, and Tudor I. Oprea. 2010 · 2010
Earlier work this paper cites.
The ddi corpus: An annotated corpus with pharmacological substances and drug–drug interactions
María Herrero-Zazo, Isabel Segura-Bedmar, Paloma Martínez, and Thierry Declerck. 2013 · 2013
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Content preserving text generation with attribute controls
Lajanugen Logeswaran, Honglak Lee, and Samy Bengio. 2018 · 2018
Earlier work this paper cites.
Genetic improvement of software: A comprehensive survey
Justyna Petke, Saemundur O. Haraldsson, Mark Harman, William B. Langdon, David R. White, and John R. Woodward. 2018 · 2018
Earlier work this paper cites.
On the summarization of consumer health questions
Asma Ben Abacha and Dina Demner-Fushman. 2019 · 2019
Earlier work this paper cites.
Transfer learning in biomedical natural language processing: An evaluation of BERT and ELMo on ten benchmarking datasets
Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
TLDR: Extreme summarization of scientific documents
Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel Weld. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
Control, generate, augment: A scalable framework for multi-attribute text generation
Giuseppe Russo, Nora Hollenstein, Claudiu Cristian Musat, and Ce Zhang. 2020 · 2020
Earlier work this paper cites.
Tweac: Transformer with extendable qa agent classifiers
Gregor Geigle, Nils Reimers, Andreas Rücklé, and Iryna Gurevych. 2021 · 2021
Earlier work this paper cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2021 · 2021
Earlier work this paper cites.
A survey of deep active learning
Pengzhen Ren, Yun Xiao, Xiaojun Chang, Po-Yao Huang, Zhihui Li, Brij B. Gupta, Xiaojiang Chen, and Xin Wang. 2021 · 2021
Earlier work this paper cites.
Attribute alignment: Controlling text generation from pre-trained language models
Dian Yu, Zhou Yu, and Kenji Sagae. 2021 · 2021
Cited alongside, same era.
Datamodels: Predicting predictions from training data
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry. 2022 · 2022
Cited alongside, same era.
Evolution through large models
Joel Lehman, Jonathan Gordon, Shawn Jain, Kamal Ndousse, Cathy Yeh, and Kenneth O. Stanley. 2022 · 2022
Cited alongside, same era.
ZeroGen: Efficient zero-shot learning via dataset generation
Jiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu, Jiangtao Feng, Zhiyong Wu, Tao Yu, and Lingpeng Kong. 2022 · 2022
Cited alongside, same era.
Central moment discrepancy (cmd) for domain-invariant representation learning
Werner Zellinger, Thomas Grubinger, Edwin Lughofer, Thomas Natschläger, and Susanne Saminger-Platz. 2022 · 2022
Cited alongside, same era.
Chain-of-interaction: Enhancing large language models for psychiatric behavior understanding by dyadic contexts
Guangzeng Han, Weisi Liu, Xiaolei Huang, and Brian Borsari. 2024 · 2024
Later among the works it cites.
Large language models as evolution strategies
Robert Lange, Yingtao Tian, and Yujin Tang. 2024 · 2024
Later among the works it cites.
Self-prompting large language models for zero-shot open-domain QA
Junlong Li, Jinyuan Wang, Zhuosheng Zhang, and Hai Zhao. 2024a · 2024
Later among the works it cites.
AutoDAN: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. 2024 · 2024
Later among the works it cites.
On LLMs-driven synthetic data generation, curation, and evaluation: A survey
Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao, Xiao Ding, Gang Chen, and Haobo Wang. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Evoprompting: Language models for code-level neural architecture search
Angelica Chen, David Dohan, and David So. 2023 · 2023
Cited alongside, same era.
Increasing diversity while maintaining accuracy: Text data generation with large language models and human interventions
John Chung, Ece Kamar, and Saleema Amershi. 2023 · 2023
Cited alongside, same era.
Exploiting asymmetry for synthetic training data generation: SynthIE and the case of information extraction
Martin Josifoski, Marija Sakota, Maxime Peyrard, and Robert West. 2023 · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Chatgpt: Optimizing language models for dialogue
OpenAI. 2022 · 2023
Cited alongside, same era.
FreeAL: Towards human-free active learning in the era of large language models
Ruixuan Xiao, Yiwen Dong, Junbo Zhao, Runze Wu, Minmin Lin, Gang Chen, and Haobo Wang. 2023 · 2023
Cited alongside, same era.
Large language model as attributed training data generator: A tale of diversity and bias
Yue Yu, Yuchen Zhuang, Jieyu Zhang, Yu Meng, Alexander J Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang. 2023 · 2023
Cited alongside, same era.
Nabeel Seedat, Nicolas Huynh, Boris van Breugel, and Mihaela van der Schaar. 2024 · 2024
Later among the works it cites.
Knowledge-infused prompting: Assessing and advancing clinical text data generation with large language models
Ran Xu, Hejie Cui, Yue Yu, Xuan Kan, Wenqi Shi, Yuchen Zhuang, May Dongmei Wang, Wei Jin, Joyce Ho, and Carl Yang. 2024 · 2024
Later among the works it cites.
Large language models as optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V Le, Denny Zhou, and Xinyun Chen. 2024 · 2024
Later among the works it cites.
Consistentchat: Building skeleton-guided consistent dialogues for large language models from scratch
Jiawei Chen, Xinyan Guan, Qianhao Yuan, Guozhao Mo, Weixiang Zhou, Yaojie Lu, Hongyu Lin, Ben He, Le Sun, and Xianpei Han. 2025 · 2025
Closest in time.
Evaluating language models as synthetic data generators
Seungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan, Seongyun Lee, Yizhong Wang, Kiril Gashteovski, Carolin Lawrence, Sean Welleck, and Graham Neubig. 2025 · 2025
Closest in time.
Contextual integrity in LLMs via reasoning and reinforcement learning
Guangchen Lan, Huseyin A Inan, Sahar Abdelnabi, Janardhan Kulkarni, Lukas Wutschitz, Reza Shokri, Christopher G Brinton, and Robert Sim. 2025 · 2025
Closest in time.
Examining and adapting time for multilingual classification via mixture of temporal experts
Weisi Liu, Guangzeng Han, and Xiaolei Huang. 2025 · 2025
Closest in time.
Multiconir: Towards multi-condition information retrieval
Xuan Lu, Sifan Liu, Bochao Yin, Yongqi Li, Xinghao Chen, Hui Su, Yaohui Jin, Wenjun Zeng, and Xiaoyu Shen. 2025 · 2025
Closest in time.
APT: Improving specialist LLM performance with weakness case acquisition and iterative preference training
Jun Rao, Zepeng Lin, Xuebo Liu, Xiaopeng Ke, Lian Lian, Dong Jin, Shengjun Cheng, Jun Yu, and Min Zhang. 2025 · 2025
Closest in time.
Mayira Sharif, Guangzeng Han, Weisi Liu, and Xiaolei Huang. 2025 · 2025
Closest in time.
Doubling your data in minutes: Ultra-fast tabular data generation via llm-induced dependency graphs
Shuo Yang, Zheyu Zhang, Bardh Prenkaj, and Gjergji Kasneci. 2025 · 2025
Closest in time.
Reso: A reward-driven self-organizing llm-based multi-agent system for reasoning tasks
Heng Zhou, Hejia Geng, Xiangyuan Xue, Li Kang, Yiran Qin, Zhiyong Wang, Zhenfei Yin, and Lei Bai. 2025 · 2025
Closest in time.