Fetching the paper…
Reading the bibliography…
Unstructured data formats account for over 80% of the data currently stored, and extracting value from such formats remains a considerable challenge.
Predicate migration: Optimizing queries with expensive predicates. In Proceedings of the 1993 ACM SIGMOD international conference on Management of data . 267–276
Joseph M Hellerstein and Michael Stonebraker. 1993 · 1993
Earlier work this paper cites.
Querying semi-structured data. In Database Theory—ICDT’97: 6th International Conference Delphi, Greece, January 8–10, 1997 Proceedings 6 . Springer, 1–18
Serge Abiteboul. 1997 · 1997
Earlier work this paper cites.
Lore: A database management system for semistructured data
Jason McHugh, Serge Abiteboul, Roy Goldman, Dallas Quass, and Jennifer Widom. 1997 · 1997
Earlier work this paper cites.
Data on the web: from relations to semistructured data and XML
Serge Abiteboul, Peter Buneman, and Dan Suciu. 2000 · 2000
Earlier work this paper cites.
Snowball: Extracting relations from large plain-text collections. In Proceedings of the fifth ACM conference on Digital libraries . 85–94
Eugene Agichtein and Luis Gravano. 2000 · 2000
Earlier work this paper cites.
Information retrieval on the web
Mei Kobayashi and Koichi Takeda. 2000 · 2000
Earlier work this paper cites.
Wrapper induction: Efficiency and expressiveness
Nicholas Kushmerick. 2000 · 2000
Earlier work this paper cites.
Modern information retrieval: A brief overview
Amit Singhal et al · 2001
Earlier work this paper cites.
Shreddr: pipelined paper digitization for low-resource organizations. In Proceedings of the 2nd ACM Symposium on Computing for Development . 1–10
Kuang Chen, Akshay Kannan, Yoriyasu Yano, Joseph M Hellerstein, and Tapan S Parikh. 2012 · 2012
Earlier work this paper cites.
NaLIR: an interactive natural language interface for querying relational databases. In Proceedings of the 2014 ACM SIGMOD international conference on Management of data . 709–712
Fei Li and Hosagrahar V Jagadish. 2014 · 2014
Earlier work this paper cites.
Extracting logical hierarchical structure of HTML documents based on headings
Tomohiro Manabe and Keishi Tajima. 2015 · 2015
Earlier work this paper cites.
https://www.forbes.com/sites/rkulkarni/2019/02/07/big-data-goes-big/?sh=45b1c73420d7
2019 · 2019
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Deep entity matching with pre-trained language models
Yuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan, and Wang-Chiew Tan. 2020 · 2020
Earlier work this paper cites.
TaBERT: Pretraining for joint understanding of textual and tabular data
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020 · 2020
Earlier work this paper cites.
https://mitsloan.mit.edu/ideas-made-to-matter/tapping-power-unstructured-data
2021 · 2021
Earlier work this paper cites.
Data provenance
Boris Glavic et al · 2021
Earlier work this paper cites.
James Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri, Sebastian Riedel, and Alon Halevy. 2021a · 2021
Earlier work this paper cites.
Text-to-table: A new way of information extraction
Xueqing Wu, Jiacheng Zhang, and Hang Li. 2021 · 2021
Earlier work this paper cites.
Recent advances in retrieval-augmented text generation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 3417–3419
Deng Cai, Yan Wang, Lemao Liu, and Shuming Shi. 2022 · 2022
Earlier work this paper cites.
Turl: Table understanding through representation learning
Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu. 2022 · 2022
Earlier work this paper cites.
DeepJoin: Joinable Table Discovery with Pre-trained Language Models
Yuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto, and Masafumi Oyamada. 2022 · 2022
Cited alongside, same era.
Few-shot learning with retrieval augmented language models
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022 · 2022
Cited alongside, same era.
A survey on retrieval-augmented text generation
Huayang Li, Yixuan Su, Deng Cai, Yan Wang, and Lemao Liu. 2022 · 2022
Cited alongside, same era.
Can foundation models wrangle your data?
Avanika Narayan, Ines Chami, Laurel Orr, Simran Arora, and Christopher Ré. 2022 · 2022
Cited alongside, same era.
Natural language interfaces to data
WannaDB: Ad-hoc SQL Queries over Text Collections. In BTW 2023 . Gesellschaft für Informatik eV, 157–181
Benjamin Hättasch, Jan-Micha Bodensohn, Liane Vogel, Matthias Urban, and Carsten Binnig. 2023 · 2023
Later among the works it cites.
Demonstration of ThalamusDB: Answering Complex SQL Queries with Natural Language Predicates on Multi-Modal Data. In Companion of the 2023 International Conference on Management of Data . 179–182
Saehan Jo and Immanuel Trummer. 2023 · 2023
Later among the works it cites.
CHORUS: foundation models for unified data discovery and exploration
Moe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou, Dan Olteanu, and Dan Suciu. 2023 · 2023
Later among the works it cites.
Revisiting prompt engineering via declarative crowdsourcing
Aditya G Parameswaran, Shreya Shankar, Parth Asawa, Naman Jain, and Yujie Wang. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abdul Quamar, Vasilis Efthymiou, Chuan Lei, Fatma Özcan, et al · 2022
Cited alongside, same era.
DB-BERT: a Database Tuning Tool that" Reads the Manual". In Proceedings of the 2022 international conference on management of data . 190–203
Immanuel Trummer. 2022 · 2022
Cited alongside, same era.
gemini.google.com
2023 · 2023
Cited alongside, same era.
https://cloud.google.com/document-ai?hl=en
2023 · 2023
Cited alongside, same era.
https://cloud.google.com/document-ai?hl=en
2023 · 2023
Cited alongside, same era.
https://openai.com/pricing
2023 · 2023
Cited alongside, same era.
https://www.anthropic.com/news/claude-3-family
2023 · 2023
Cited alongside, same era.
https://www.forbes.com/sites/stevemcdowell/2023/03/09/komprise-unleashes-fresh-insights-about-your-unstructured-data/?sh=5f444c474aa9
2023 · 2023
Cited alongside, same era.
Mohammed Saeed, Nicola De Cao, and Paolo Papotti. 2023 · 2023
Later among the works it cites.
Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?
Xiangru Tang, Yiming Zong, Yilun Zhao, Arman Cohan, and Mark Gerstein. 2023 · 2023
Later among the works it cites.
Towards Multi-Modal DBMSs for Seamless Querying of Texts and Tables
Matthias Urban and Carsten Binnig. 2023 · 2023
Later among the works it cites.
OmniscientDB: a large language model-augmented DBMS that knows what other DBMSs do not know. In Proceedings of the Sixth International Workshop on Exploiting Artificial Intelligence Techniques for Data Management . 1–7
Matthias Urban, Duc Dat Nguyen, and Carsten Binnig. 2023 · 2023
Later among the works it cites.
Learning to Filter Context for Retrieval-Augmented Generation
Zhiruo Wang, Jun Araki, Zhengbao Jiang, Md Rizwan Parvez, and Graham Neubig. 2023 · 2023
Later among the works it cites.
Large language models as data preprocessors
Haochen Zhang, Yuyang Dong, Chuan Xiao, and Masafumi Oyamada. 2023 · 2023
Later among the works it cites.
http://personal-informatics.depstein.net
2024 · 2024
Closest in time.
https://primis.phmsa.dot.gov/enforcement-data/cases/NOPV
2024 · 2024
Closest in time.
https://pypi.org/project/pdfplumber/0.1.2/
2024 · 2024
Closest in time.
https://www.malibucity.org/AgendaCenter
2024 · 2024
Closest in time.
Large Language Models on Tabular Data–A Survey
Xi Fang, Weijie Xu, Fiona Anting Tan, Jiani Zhang, Ziqing Hu, Yanjun Qi, Scott Nickleach, Diego Socolinsky, Srinivasan Sengamedu, and Christos Faloutsos. 2024 · 2024
Closest in time.
Cocoon: Semantic Table Profiling Using Large Language Models
Zezhou Huang and Eugene Wu. 2024 · 2024
Closest in time.
Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls
Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, et al · 2024
Closest in time.
Query Rewriting via Large Language Models
Jie Liu and Barzan Mozafari. 2024 · 2024
Closest in time.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024 · 2024
Closest in time.
RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval. In The Twelfth International Conference on Learning Representations
Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna, Anna Goldie, and Christopher D Manning. 2024 · 2024
Closest in time.