Fetching the paper…
Reading the bibliography…
Recent smaller language models such Phi-3.5 and Phi-4 rely on synthetic data generated using larger Language models.
Task2vec: Task embedding for meta-learning
Alessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran, Subhransu Maji, Charless Fowlkes, Stefano Soatto, and Pietro Perona. 2019 · 1902
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Www’18 open challenge: Financial opinion mining and question answering
Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018 · 1942
Earlier work this paper cites.
Using tf-idf to determine word relevance in document queries
Juan Enrique Ramos. 2003 · 2003
Earlier work this paper cites.
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2020 · 2009
Earlier work this paper cites.
Impact of news on the commodity market: Dataset and results
Ankur Sinha and Tanmay Khandait. 2020 · 2009
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
Pekka Malo, Ankur Sinha, Pyry Takala, Pekka Korhonen, and Jyrki Wallenius. 2013 · 2013
Earlier work this paper cites.
Domain adaption of named entity recognition to support credit risk assessment
Julio Cesar Salinas Alvarado, Karin Verspoor, and Timothy Baldwin. 2015 · 2015
Earlier work this paper cites.
Chemprot-3.0: a global chemical biology diseases mapping
Jens Vindahl Kringelum, Sonny Kim Kjærulff, Søren Brunak, Ole Lund, Tudor I. Oprea, and Olivier Taboureau. 2016 · 2016
Earlier work this paper cites.
PubMed 200k RCT: a dataset for sequential sentence classification in medical abstracts
Franck Dernoncourt and Ji Young Lee. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Effective transfer learning for identifying similar questions: Matching user questions to covid-19 faqs
Clara H. McCreery, Namit Katariya, Anitha Kannan, Manish Chablani, and Xavier Amatriain. 2020 · 2020
Earlier work this paper cites.
Directed diversity: Leveraging language embedding distances for collective creativity in crowd ideation
Samuel Rhys Cox, Yunlong Wang, Ashraf Abdul, Christian von der Weth, and Brian Y. Lim. 2021 · 2021
Earlier work this paper cites.
Convfinqa: Exploring the chain of numerical reasoning in conversational finance question answering
Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang. 2022 · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. 2022 · 2022
Earlier work this paper cites.
Unnatural instructions: Tuning language models with (almost) no human labor
Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2022 · 2022
Earlier work this paper cites.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Maitreya Patel, Kuntal Kumar Pal, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Shailaja Keyur Sampat, Savan Doshi, Siddhartha Mishra, Sujan Reddy, Sumanta Patro, Tanay Dixit, Xudong Shen, Chitta Baral, Yejin Choi, Noah A. Smith, Hannaneh Hajishirzi, and Daniel Khashabi. 2022 · 2022
Earlier work this paper cites.
ADEQA: A question answer based approach for joint ADE-suspect extraction using sequence-to-sequence transformers
Vinayak Arannil, Tomal Deb, and Atanu Roy. 2023 · 2023
Earlier work this paper cites.
Increasing diversity while maintaining accuracy: Text data generation with large language models and human interventions
John Chung, Ece Kamar, and Saleema Amershi. 2023 · 2023
Cited alongside, same era.
Enhancing chat language models by scaling high-quality instructional conversations
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023 · 2023
Cited alongside, same era.
Tinystories: How small can language models be and still speak coherent english?
Ronen Eldan and Yuanzhi Li. 2023 · 2023
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Cited alongside, same era.
Aaron Grattafiori et al. 2024 · 2024
Later among the works it cites.
Large language model based multi-agents: A survey of progress and challenges
Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V. Chawla, Olaf Wiest, and Xiangliang Zhang. 2024 · 2024
Later among the works it cites.
How far can we extract diverse perspectives from large language models?
Shirley Anugrah Hayati, Minhwa Lee, Dheeraj Rajagopal, and Dongyeop Kang. 2024 · 2024
Later among the works it cites.
Evaluating language models as synthetic data generators
Seungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan, Seongyun Lee, Yizhong Wang, Kiril Gashteovski, Carolin Lawrence, Sean Welleck, and Graham Neubig. 2024 · 2024
Later among the works it cites.
Self-prompting large language models for zero-shot open-domain QA
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alycia Lee, Brando Miranda, and Sanmi Koyejo. 2023 · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023 · 2023
Cited alongside, same era.
Zhuoyan Li, Hangxiao Zhu, Zhuoran Lu, and Ming Yin. 2023 · 2023
Cited alongside, same era.
Explore-instruct: Enhancing domain-specific instruction coverage through active exploration
Fanqi Wan, Xinting Huang, Tao Yang, Xiaojun Quan, Wei Bi, and Shuming Shi. 2023 · 2023
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023 · 2023
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang. 2023 · 2023
Cited alongside, same era.
Wizardlm: Empowering large language models to follow complex instructions
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023 · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023 · 2023
Cited alongside, same era.
Junlong Li, Jinyuan Wang, Zhuosheng Zhang, and Hai Zhao. 2024 · 2024
Later among the works it cites.
Agentinstruct: Toward generative teaching with agentic flows
Arindam Mitra, Luciano Del Corro, Guoqing Zheng, Shweti Mahajan, Dany Rouhana, Andres Codas, Yadong Lu, Wei ge Chen, Olga Vrousgos, Corby Rosset, Fillipe Silva, Hamed Khanpour, Yash Lara, and Ahmed Awadallah. 2024 · 2024
Later among the works it cites.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2024 · 2024
Later among the works it cites.
How bad is training on synthetic data? a statistical analysis of language model collapse
Mohamed El Amine Seddik, Suei-Wen Chen, Soufiane Hayou, Pierre Youssef, and Merouane Debbah. 2024 · 2024
Later among the works it cites.
The curse of recursion: Training on generated data makes models forget
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson. 2024 · 2024
Later among the works it cites.
Meta-prompting: Enhancing language models with task-agnostic scaffolding
Mirac Suzgun and Adam Tauman Kalai. 2024 · 2024
Later among the works it cites.
Will we run out of data? limits of llm scaling based on human-generated data
Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, and Marius Hobbhahn. 2024 · 2024
Later among the works it cites.
Simplestrat: Diversifying language model generation with stratification
Justin Wong, Yury Orlovskiy, Michael Luo, Sanjit A. Seshia, and Joseph E. Gonzalez. 2024 · 2024
Later among the works it cites.
Unigen: A unified framework for textual dataset generation using large language models
Siyuan Wu, Yue Huang, Chujie Gao, Dongping Chen, Qihui Zhang, Yao Wan, Tianyi Zhou, Xiangliang Zhang, Jianfeng Gao, Chaowei Xiao, and Lichao Sun. 2024 · 2024
Later among the works it cites.
Yifan Zhang, Yang Yuan, and Andrew Chi-Chih Yao. 2024 · 2024
Later among the works it cites.
Introducing the next generation of claude
Anthropic. 2024 · 2025
Closest in time.
Demystifying domain-adaptive post-training for financial llms
Zixuan Ke, Yifei Ming, Xuan-Phi Nguyen, Caiming Xiong, and Shafiq Joty. 2025 · 2025
Closest in time.
Concise thoughts: Impact of output length on llm reasoning and cost
Sania Nayab, Giulio Rossolini, Marco Simoni, Andrea Saracino, Giorgio Buttazzo, Nicolamaria Manes, and Fabrizio Giacomelli. 2025 · 2025
Closest in time.
Kaiser Sun and Mark Dredze. 2025 · 2025
Closest in time.
Synthesizing post-training data for llms through multi-agent simulation
Shuo Tang, Xianghe Pang, Zexi Liu, Bohan Tang, Rui Ye, Tian Jin, Xiaowen Dong, Yanfeng Wang, and Siheng Chen. 2025 · 2025
Closest in time.
Ran Xu, Hejie Cui, Yue Yu, Xuan Kan, Wenqi Shi, Yuchen Zhuang, Wei Jin, Joyce Ho, and Carl Yang. 2025 · 2025
Closest in time.