Fetching the paper…
Reading the bibliography…
A long standing goal of the data management community is to develop general, automated systems that ingest semi-structured documents and output queryable tables without human effort or domain specific customization.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Extracting patterns and relations from the WorldWide Web. In WebDB
S. Brin. 1998 · 1998
Earlier work this paper cites.
XTRACT: A system for extracting document type descriptors from XML documents. In Proceedings of the 2000 ACM SIGMOD international conference on Management of data . 165–176
Minos Garofalakis, Aristides Gionis, Rajeev Rastogi, Sridhar Seshadri, and Kyuseok Shim. 2000 · 2000
Earlier work this paper cites.
Snowball: Extracting Relations from Large Plain-Text Collections. In DL ’00: Proceedings of the fifth ACM conference on Digital libraries
Eugene Agichtein Luis Gravano. 2000 · 2000
Earlier work this paper cites.
Information extraction and integration: An overview
W. Cohen. 2004 · 2004
Earlier work this paper cites.
Introducing the enron corpus. In Proceedings of the 1st Conference on Email and Anti-Spam (CEAS)
B. Klimt and Y. Yang. 2004 · 2004
Earlier work this paper cites.
Avatar information extraction system
S. Raghavan S. Vaithyanathan T.S. Jayram, R. Krishnamurthy and H. Zhu. 2006 · 2006
Earlier work this paper cites.
Open information extraction from the web
Michele Banko, Michael J. Cafarella, Stephen Soderland, Matthew G Broadhead, and Oren Etzioni. 2007 · 2007
Earlier work this paper cites.
A Relational Approach to Incrementally Extracting and Querying Structure in Unstructured Data. In VLDB
Eric Chu, Akanksha Baid, Ting Chen, AnHai Doan, and Jeffrey Naughton. 2007 · 2007
Earlier work this paper cites.
Building structured web community portals: A top-down, compositional, and incremental approach
F. Chen A. Doan P. DeRose, W. Shen and R. Ramakrishnan. 2007 · 2007
Earlier work this paper cites.
Open information extraction from the web
Oren Etzioni, Michele Banko, Stephen Soderland, and Daniel S Weld. 2008 · 2008
Earlier work this paper cites.
From one tree to a forest: a unified solution for structured web data extraction
Qiang Hao, Rui Cai, Yanwei Pang, and Lei Zhang. 2011 · 2011
Earlier work this paper cites.
Medical Device Recalls and the FDA Approval Process
Diana M. Zuckerman, Paul Brown, and Steven E. Nissen. 2011 · 2011
Earlier work this paper cites.
Open Language Learning for Information Extraction
Mausam, Michael Schmitz, Stephen Soderland, Robert Bart, and Oren Etzioni. 2012 · 2012
Earlier work this paper cites.
Extraction and integration of partially overlapping web sources
Mirko Bronzi, Valter Crescenzi, Paolo Merialdo, and Paolo Papotti. 2013 · 2013
Earlier work this paper cites.
Data mining in education. In Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery
C. Romero and S. Ventura. 2013 · 2013
Earlier work this paper cites.
Managing unstructured data in relational databases. In 2013 IEEE Conference on Systems, Process & Control (ICSPC)
Wael M.S. Yafooz, Siti Z.Z. Abidin, Nasiroh Omar, and Zanariah Idrus. 2013 · 2013
Earlier work this paper cites.
A big data guide to understanding climate change: The case for theory-guided data science. In Big data
J. H. Faghmous and V Kumar. 2014 · 2014
Earlier work this paper cites.
Incremental knowledge base construction using deepdive. In Proceedings of the VLDB Endowment International Conference on Very Large Data Bases (VLDB)
Jaeho Shin, Sen Wu, Feiran Wang, Christopher De Sa, Ce Zhang, and Christopher Ré. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
What the Enron E-mails Say About Us
Nathan Heller. 2017 · 2017
Earlier work this paper cites.
Snorkel: Rapid Training Data Creation with Weak Supervision
Alexander Ratner, Stephen H. Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher Ré . 2017 · 2017
Earlier work this paper cites.
A Survey on Open Information Extraction. In Proceedings of the 27th International Conference on Computational Linguistics
Christina Niklaus, Matthias Cetto, André Freitas, and Siegfried Handschuh. 2018 · 2018
Cited alongside, same era.
Snuba: Automating Weak Supervision to Label Training Data
Paroma Varma and Christopher Ré. 2018 · 2018
Cited alongside, same era.
Reproducible, interactive, scalable and extensible microbiome data science using qiime 2. In Nature biotechnology
Rideout J. R. Dillon M. R. Bokulich N. A. Abnet C. C. Al-Ghalith G. A. Alexander H. Alm E. J. Arumugam M. et al. Bolyen, E. 2019 · 2019
Cited alongside, same era.
OpenCeres: When Open Information Extraction Meets the Semi-Structured Web
Colin Lockard, Prashant Shiralkar, and Xin Luna Dong. 2019 · 2019
Cited alongside, same era.
Data Lake Management: Challenges and Opportunities
Fatemeh Nargesian, Erkang Zhu, Reneé J. Miller, Ken Q. Pu, and Patricia C. Arocena. 2019 · 2019
Cited alongside, same era.
A Survey on Retrieval-Augmented Text Generation
Huayang Li, Yixuan Su, Deng Cai, Yan Wang, and Lemao Liu. 2022 · 2022
Later among the works it cites.
Holistic Evaluation of Language Models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, and more. 2022 · 2022
Later among the works it cites.
PEFT: State-of-the-art Parameter-Efficient Fine-Tuning methods
Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, and Sayak Paul. 2022 · 2022
Later among the works it cites.
Can Foundation Models Wrangle Your Data?
Avanika Narayan, Ines Chami, Laurel Orr, Simran Arora, and Christopher Ré. 2022 · 2022
Later among the works it cites.
Manifest
Laurel Orr. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning Dependency Structures for Weak Supervision Models
Paroma Varma, Frederic Sala, Ann He, Alexander Ratner, and Christopher Re. 2019 · 2019
Cited alongside, same era.
ELECTRA: pre-training text encoders as discriminators rather than generators. In International Conference on Learning Representations (ICLR)
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Cited alongside, same era.
Fast and Three-rious: Speeding Up Weak Supervision with Triplet Methods. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research) , Vol. 119. PMLR, 3280–3291
Daniel Fu, Mayee Chen, Frederic Sala, Sarah Hooper, Kayvon Fatahalian, and Christopher Re. 2020 · 2020
Cited alongside, same era.
OpenIE6: Iterative Grid Labeling and Coordination Analysis for Open Information Extraction. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)
Keshav Kolluru, Vaibhav Adlakha, Samarth Aggarwal, Mausam, and Soumen Chakrabarti. 2020 · 2020
Cited alongside, same era.
ZeroShotCeres: Zero-Shot Relation Extraction from Semi-Structured Webpages
Colin Lockard, Prashant Shiralkar, Xin Luna Dong, and Hannaneh Hajishirzi. 2020 · 2020
Cited alongside, same era.
A General Language Assistant as a Laboratory for Alignment
Amanda Askell, Yushi Bai, Anna Chen, Dawn Drain, Deep Ganguli, T. J. Henighan, Andy Jones, and Nicholas Joseph et al. 2021 · 2021
Cited alongside, same era.
Interactive weak supervision: Learning useful heuristics for data labeling
Benedikt Boecking, Willie Neiswanger, Eric Xing, and Artur Dubrawski. 2021 · 2021
Cited alongside, same era.
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Operationalizing Machine Learning: An Interview Study
Shreya Shankar, Rolando Garcia, Joseph M. Hellerstein, and Aditya G. Parameswaran. 2022 · 2022
Later among the works it cites.
Language Models in the Loop: Incorporating Prompting into Weak Supervision
Ryan Smith, Jason A. Fries, Braden Hancock, and Stephen H. Bach. 2022 · 2022
Later among the works it cites.
CodexDB: synthesizing code for query processing from natural language instructions using GPT-3 codex
Immanuel Trummer. 2022 · 2022
Later among the works it cites.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le Le, Ed H. Cho, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022 · 2022
Later among the works it cites.
Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts. In CHI Conference on Human Factors in Computing Systems . 1–22
Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022 · 2022
Later among the works it cites.
A Survey on Neural Open Information Extraction: Current Status and Future Directions
Shaowen Zhou, Bowen Yu, Aixin Sun, Cheng Long, Jingyang Li, Haiyang Yu, Jian Sun, and Yongbin Li. 2022 · 2022
Later among the works it cites.
Wikipedia Statistics
April 2023 · 2023
Closest in time.
Reasoning over Public and Private Data in Retrieval-Based Systems
Simran Arora, Patrick Lewis, Angela Fan, Jacob Kahn, and Christopher Ré. 2023a · 2023
Closest in time.
Ask Me Anything: A simple strategy for prompting language models
Simran Arora, Avanika Narayan, Mayee F. Chen, Laurel Orr, Neel Guha, Kush Bhatia, Ines Chami, Frederic Sala, and Christopher Ré. 2023b · 2023
Closest in time.
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel Trummer, and Christopher Ré. 2023c · 2023
Closest in time.
The Safety of Inpatient Health Care
David W Bates, David M Levine, Hojjat Salmasian, Ania Syrowatka, David M Shahian, Stuart Lipsitz, Jonathan P Zebrowski, Laura C Myers, Merranda S Logan, Christopher G Roy, et al · 2023
Closest in time.
Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes
Zui Chen, Zihui Gu, Lei Cao, Ju Fan, Sam Madden, and Nan Tang. 2023 · 2023
Closest in time.
How Many Websites Are There in the World?
Nick Huss. 2023 · 2023
Closest in time.
ChatGPT: Jack of all trades, master of none
Jan Kocoń, Igor Cichecki, Oliwier Kaszyca, Mateusz Kochanek, Dominika Szydło, Joanna Baran, Julita Bielaniewicz, Marcin Gruza, Arkadiusz Janz, Kamil Kanclerz, et al · 2023
Closest in time.
OpenAI API
OpenAI. March 2023 · 2023
Closest in time.
High-throughput Generative Inference of Large Language Models with a Single GPU
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Daniel Y Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E Gonzalez, et al · 2023
Closest in time.