Fetching the paper…
Reading the bibliography…
Large Language models (LLMs) usually rely on extensive training datasets.
W. Wang, T. Hao, and W. Liu, “Automatic question generation for learning evaluation in medicine,” in Advances in Web Based Learning-ICWL International Conference , vol. 4823. Springer Science & Business Media, 2008, p. 242
2008
Earlier work this paper cites.
B. E. Needles, M. Powers, and S. V. Crosson, Principles of accounting . Cengage Learning, 2013
2013
Earlier work this paper cites.
G. Tsatsaronis, G. Balikas, P. Malakasiotis, I. Partalas, M. Zschunke, M. R. Alvers, D. Weissenborn, A. Krithara, S. Petridis, D. Polychronopoulos, Y. Almirantis, J. Pavlopoulos, N. Baskiotis, P. Gallinari, T. Artières, A. N. Ngomo, N. Heino, É. Gaussier, L. Barrio-Alvers, M. Schroeder, I. Androutsopoulos, and G. Paliouras, “An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition,” BMC Bioinform. , vol. 16, pp. 138:1–138:28, 2015
2015
Earlier work this paper cites.
J. Araki, D. Rajagopal, S. Sankaranarayanan, S. Holm, Y. Yamakawa, and T. Mitamura, “Generating questions and multiple-choice answers using semantic analysis of texts,” in International Conference on Computational Linguistics , 2016, pp. 1125–1136
2016
Earlier work this paper cites.
C. Song, T. Ristenpart, and V. Shmatikov, “Machine learning models that remember too much,” in ACM SIGSAC Conference on computer and communications security , 2017, pp. 587–601
2017
Earlier work this paper cites.
J. Han, U. Barman, J. Hayes, J. Du, E. Burgin, and D. Wan, “Nextgen AML: distributed deep learning based language technologies to augment anti money laundering investigation,” in Conference of the Association for Computational Linguistics . Association for Computational Linguistics, 2018, pp. 37–42
2018
Earlier work this paper cites.
N. Park, M. Mohammadi, K. Gorde, S. Jajodia, H. Park, and Y. Kim, “Data synthesis based on generative adversarial networks,” VLDB Endowment , vol. 11, no. 10, pp. 1071–1083, 2018
2018
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner, “DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs,” in Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . Association for Computational Linguistics, 2019, pp. 2368–2378
2019
Earlier work this paper cites.
C. Alberti, D. Andor, E. Pitler, J. Devlin, and M. Collins, “Synthetic QA corpora generation with roundtrip consistency,” in Conference of the Association for Computational Linguistics . Association for Computational Linguistics, 2019, pp. 6168–6173
2019
Earlier work this paper cites.
L. Xu, M. Skoularidou, A. Cuesta-Infante, and K. Veeramachaneni, “Modeling tabular data using conditional GAN,” in Conference on Neural Information Processing Systems , 2019, pp. 7333–7343
2019
Earlier work this paper cites.
S. Li, B. Tai, and Y. Huang, “Evaluating variational autoencoder as a private data release mechanism for tabular data,” in IEEE Pacific Rim International Symposium on Dependable Computing , 2019, pp. 198–206
2019
Earlier work this paper cites.
W. Chen, H. Wang, J. Chen, Y. Zhang, H. Wang, S. Li, X. Zhou, and W. Y. Wang, “Tabfact: A large-scale dataset for table-based fact verification,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
W. Chen, H. Zha, Z. Chen, W. Xiong, H. Wang, and W. Y. Wang, “Hybridqa: A dataset of multi-hop question answering over tabular and textual data,” in Findings of the Association for Computational Linguistics: EMNLP . Association for Computational Linguistics, 2020, pp. 1026–1036
2020
Cited alongside, same era.
S. Shen, Y. Li, N. Du, X. Wu, Y. Xie, S. Ge, T. Yang, K. Wang, X. Liang, and W. Fan, “On the generation of medical question-answer pairs,” in AAAI Conference on Artificial Intelligence , vol. 34, no. 05, 2020, pp. 8822–8829
2020
Cited alongside, same era.
S. A. Assefa, D. Dervovic, M. Mahfouz, R. E. Tillman, P. Reddy, and M. Veloso, “Generating synthetic data in finance: opportunities, challenges and pitfalls,” in ACM International Conference on AI in Finance , 2020, pp. 1–8
2020
Cited alongside, same era.
Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, and W. Y. Wang, “FinQA: A dataset of numerical reasoning over financial data,” in Conference on Empirical Methods in Natural Language Processing , 2021, pp. 3697–3711
Z. Chen, S. Li, C. Smiley, Z. Ma, S. Shah, and W. Y. Wang, “ConvFinQA: Exploring the chain of numerical reasoning in conversational finance question answering,” in Conference on Empirical Methods in Natural Language Processing , 2022, pp. 6279–6292
2022
Later among the works it cites.
H. Brown, K. Lee, F. Mireshghallah, R. Shokri, and F. Tramèr, “What does it mean for a language model to preserve privacy?” in ACM Conference on Fairness, Accountability, and Transparency , 2022, pp. 2280–2292
2022
Later among the works it cites.
2023
Later among the works it cites.
X. Li, Y. Zhu, S. Liu, J. Ju, Y. Qu, and G. Cheng, “Dyrren: A dynamic retriever-reranker-generator model for numerical reasoning over tabular and textual data,” AAAI Conference on Artificial Intelligence , vol. 37, no. 11, pp. 13 139–13 147, Jun. 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
F. Zhu, W. Lei, Y. Huang, C. Wang, S. Zhang, J. Lv, F. Feng, and T.-S. Chua, “TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance,” in Annual Meeting of the Association for Computational Linguistics , 2021, pp. 3277–3287
2021
Cited alongside, same era.
C. Lyu, L. Shang, Y. Graham, J. Foster, X. Jiang, and Q. Liu, “Improving unsupervised question answering via summarization-informed question generation,” in Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 2021, pp. 4134–4148
2021
Cited alongside, same era.
Z. Zhao, A. Kunar, R. Birke, and L. Y. Chen, “CTAB-GAN: effective table data synthesizing,” in Asian Conference on Machine Learning . PMLR, 2021, pp. 97–112
2021
Cited alongside, same era.
2021
Cited alongside, same era.
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson et al. , “Extracting training data from large language models,” in USENIX Security Symposium , 2021, pp. 2633–2650
2021
Cited alongside, same era.
F. Lei, S. He, X. Li, J. Zhao, and K. Liu, “Answering numerical reasoning questions in table-text hybrid contents with graph-based encoder and tree-based decoder,” in International Conference on Computational Linguistics , 2022, pp. 1379–1390
2022
Cited alongside, same era.
Y. Zhao, Y. Li, C. Li, and R. Zhang, “Multihiertt: Numerical reasoning over multi hierarchical tabular and textual data,” in Annual Meeting of the Association for Computational Linguistics , 2022, pp. 6588–6600
2022
Cited alongside, same era.
Y. Zhao, Y. Li, C. Li, and R. Zhang, “Multihiertt: Numerical reasoning over multi hierarchical tabular and textual data,” in Annual Meeting of the Association for Computational Linguistics , 2022, pp. 6588–6600
2022
Cited alongside, same era.
2023
Later among the works it cites.
S. Lee, H. Kim, and J. Kang, “LIQUID: A framework for list question answering dataset generation,” in AAAI Conference on Artificial Intelligence , 2023, pp. 13 014–13 024
2023
Later among the works it cites.
A. Rogers, M. Gardner, and I. Augenstein, “QA dataset explosion: A taxonomy of NLP resources for question answering and reading comprehension,” ACM Computing Surveys , vol. 55, no. 10, pp. 197:1–197:45, 2023
2023
Later among the works it cites.
V. Borisov, K. Seßler, T. Leemann, M. Pawelczyk, and G. Kasneci, “Language models are realistic tabular data generators,” in International Conference on Learning Representations , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Kotelnikov, D. Baranchuk, I. Rubachev, and A. Babenko, “Tabddpm: Modelling tabular data with diffusion models,” in International Conference on Machine Learning , vol. 202. PMLR, 2023, pp. 17 564–17 579
2023
Later among the works it cites.
F. Zhu, M. Li, J. Xiao, F. Feng, C. Wang, and T. S. Chua, “Soargraph: Numerical reasoning over financial table-text data via semantic-oriented hierarchical graphs,” in Companion Proceedings of the ACM Web Conference 2023 , 2023, pp. 1236–1244
2023
Later among the works it cites.
2023
Later among the works it cites.
I. Shumailov, Z. Shumaylov, Y. Zhao, Y. Gal, N. Papernot, and R. Anderson, “Model dementia: Generated data makes models forget,” arXiv e-prints , pp. arXiv–2305, 2023
2023
Later among the works it cites.