Fetching the paper…
Reading the bibliography…
Analyzing unstructured data has been a persistent challenge in data processing.
CDB: A crowd-powered database system
Guoliang Li, Chengliang Chai, Ju Fan, Xueping Weng, Jian Li, Yudian Zheng, Yuanbing Li, Xiang Yu, Xiaohang Zhang, and Haitao Yuan. 2018 · 1929
Earlier work this paper cites.
Access path selection in a relational database management system. In Proceedings of the 1979 ACM SIGMOD international conference on Management of data . 23–34
P Griffiths Selinger, Morton M Astrahan, Donald D Chamberlin, Raymond A Lorie, and Thomas G Price. 1979 · 1979
Earlier work this paper cites.
Maintaining views incrementally
Ashish Gupta, Inderpal Singh Mumick, and Venkatramanan Siva Subrahmanian. 1993 · 1993
Earlier work this paper cites.
The Cascades Framework for Query Optimization
Goetz Graefe. 1995 · 1995
Earlier work this paper cites.
An overview of query optimization in relational systems. In Proceedings of the seventeenth ACM SIGACT-SIGMOD-SIGART symposium on Principles of database systems . 34–43
Surajit Chaudhuri. 1998 · 1998
Earlier work this paper cites.
Anatomy of a database system
Joseph M Hellerstein and Michael Stonebraker. 2005 · 2005
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
MapReduce online.. In Nsdi , Vol. 10. 20
Tyson Condie, Neil Conway, Peter Alvaro, Joseph M Hellerstein, Khaled Elmeleegy, and Russell Sears. 2010 · 2010
Earlier work this paper cites.
CrowdDB: answering queries with crowdsourcing. In Proceedings of the 2011 ACM SIGMOD International Conference on Management of data . 61–72
Michael J Franklin, Donald Kossmann, Tim Kraska, Sukriti Ramesh, and Reynold Xin. 2011 · 2011
Earlier work this paper cites.
Crowdsourced databases: Query processing with people. Cidr
Adam Marcus, Eugene Wu, David R Karger, Samuel Madden, and Robert C Miller. 2011 · 2011
Earlier work this paper cites.
Deco: declarative crowdsourcing. In Proceedings of the 21st ACM international conference on Information and knowledge management . 1203–1212
Aditya Ganesh Parameswaran, Hyunjung Park, Hector Garcia-Molina, Neoklis Polyzotis, and Jennifer Widom. 2012 · 2012
Earlier work this paper cites.
Vader: A parsimonious rule-based model for sentiment analysis of social media text. In Proceedings of the international AAAI conference on web and social media , Vol. 8. 216–225
Clayton Hutto and Eric Gilbert. 2014 · 2014
Earlier work this paper cites.
Idk cascades: Fast deep learning by learning not to overthink
Xin Wang, Yujia Luo, Daniel Crankshaw, Alexey Tumanov, Fisher Yu, and Joseph E Gonzalez. 2017 · 2017
Earlier work this paper cites.
An Overview of End-to-End Entity Resolution for Big Data
Vassilis Christophides, Vasilis Efthymiou, Themis Palpanas, George Papadakis, and Kostas Stefanidis. 2020 · 2020
Earlier work this paper cites.
spaCy: Industrial-strength Natural Language Processing in Python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020 · 2020
Earlier work this paper cites.
CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review
Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball. 2021 · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Earlier work this paper cites.
DeepJoin: Joinable Table Discovery with Pre-trained Language Models
Yuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto, and Masafumi Oyamada. 2022 · 2022
Earlier work this paper cites.
DB-BERT: a Database Tuning Tool that" Reads the Manual". In Proceedings of the 2022 international conference on management of data . 190–203
Immanuel Trummer. 2022 · 2022
Cited alongside, same era.
Language models enable simple systems for generating structured views of heterogeneous data lakes
Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel Trummer, and Christopher Ré. 2023 · 2023
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al · 2023
Cited alongside, same era.
Observatory: Characterizing Embeddings of Relational Tables
Tianji Cong, Madelon Hulsebos, Zhenjie Sun, Paul Groth, and HV Jagadish. 2023 · 2023
Cited alongside, same era.
How Large Language Models Will Disrupt Data Management
DSPy: Compiling Declarative Language Model Calls into State-of-the-Art Pipelines. In The Twelfth International Conference on Learning Representations
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Saiful Haq, Ashutosh Sharma, Thomas T Joshi, Hanna Moazam, Heather Miller, et al · 2024
Closest in time.
Mosh Levy, Alon Jacoby, and Yoav Goldberg. 2024 · 2024
Closest in time.
Towards Accurate and Efficient Document Analytics with Large Language Models
Yiming Lin, Madelon Hulsebos, Ruiying Ma, Shreya Shankar, Sepanta Zeigham, Aditya G Parameswaran, and Eugene Wu. 2024 · 2024
Closest in time.
A Declarative System for Optimizing AI Workloads
Chunwei Liu, Matthew Russo, Michael Cafarella, Lei Cao, Peter Baille Chen, Zui Chen, Michael Franklin, Tim Kraska, Samuel Madden, and Gerardo Vitagliano. 2024b · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raul Castro Fernandez, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan, and Chenhao Tan. 2023a · 2023
Cited alongside, same era.
How large language models will disrupt data management
Raul Castro Fernandez, Aaron J Elmore, Michael J Franklin, Sanjay Krishnan, and Chenhao Tan. 2023b · 2023
Cited alongside, same era.
Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression
Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023 · 2023
Cited alongside, same era.
Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning . PMLR, 31210–31227
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Cited alongside, same era.
Found in the middle: Permutation self-consistency improves listwise ranking in large language models
Raphael Tang, Xinyu Zhang, Xueguang Ma, Jimmy Lin, and Ferhan Ture. 2023 · 2023
Cited alongside, same era.
A prompt pattern catalog to enhance prompt engineering with chatgpt
Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. 2023 · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Cited alongside, same era.
The Design of an LLM-powered Unstructured Analytics System
Eric Anderson, Jonathan Fritz, Austin Lee, Bohou Li, Mark Lindblad, Henry Lindeman, Alex Meyer, Parth Parmar, Tanvi Ranade, Mehul A. Shah, Benjamin Sowell, Dan Tecuci, Vinayak Thapliyal, and Matt Welsh. 2024 · 2024
Cited alongside, same era.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024a · 2024
Closest in time.
Revisiting Prompt Engineering via Declarative Crowdsourcing
Aditya G Parameswaran, Shreya Shankar, Parth Asawa, Naman Jain, and Yujie Wang. 2024 · 2024
Closest in time.
Semantic Operators: A Declarative Model for Rich, AI-based Analytics Over Text Data
Liana Patel, Siddharth Jha, Parth Asawa, Melissa Pan, Carlos Guestrin, and Matei Zaharia. 2024 · 2024
Closest in time.
On limitations of the transformer architecture
Binghui Peng, Srini Narayanan, and Christos Papadimitriou. 2024 · 2024
Closest in time.
Chase-sql: Multi-path reasoning and preference optimized candidate selection in text-to-sql
Mohammadreza Pourreza, Hailong Li, Ruoxi Sun, Yeounoh Chung, Shayan Talaei, Gaurav Tarlok Kakkar, Yu Gan, Amin Saberi, Fatma Ozcan, and Sercan O Arik. 2024 · 2024
Closest in time.
CleanAgent: Automating Data Standardization with LLM-based Agents
Danrui Qi and Jiannan Wang. 2024 · 2024
Closest in time.
Building Reactive Large Language Model Pipelines with Motion. In Companion of the 2024 International Conference on Management of Data . 520–523
Shreya Shankar and Aditya G Parameswaran. 2024 · 2024
Closest in time.
Confabulation: The Surprising Value of Large Language Model Hallucinations
Peiqi Sui, Eamon Duede, Sophie Wu, and Richard Jean So. 2024 · 2024
Closest in time.
Demonstrating CAESURA: Language Models as Multi-Modal Query Planners. In Companion of the 2024 International Conference on Management of Data . 472–475
Matthias Urban and Carsten Binnig. 2024 · 2024
Closest in time.
A Field Guide to Automatic Evaluation of LLM-Generated Summaries. In Annual International ACM SIGIR Conference on Research and Development in Information Retrieval
Tempest A. van Schaik and Brittany Pugh. 2024 · 2024
Closest in time.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2024 · 2024
Closest in time.
LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration
Jun Zhao, Can Zu, Hao Xu, Yi Lu, Wei He, Yiwen Ding, Tao Gui, Qi Zhang, and Xuanjing Huang. 2024 · 2024
Closest in time.
JudgeLM: Fine-tuned Large Language Models are Scalable Judges. In The Thirteenth International Conference on Learning Representations
Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2025 · 2025
Closest in time.