Fetching the paper…
Reading the bibliography…
In many use-cases, information is stored in text but not available in structured data.
CrowdDB: Answering Queries with Crowdsourcing. In ACM SIGMOD
Michael J. Franklin, Donald Kossmann, Tim Kraska, Sukriti Ramesh, and Reynold Xin. 2011 · 2011
Earlier work this paper cites.
The bigdawg polystore system
Jennie Duggan, Aaron J Elmore, Michael Stonebraker, Magda Balazinska, Bill Howe, Jeremy Kepner, Sam Madden, David Maier, Tim Mattson, and Stan Zdonik. 2015 · 2015
Earlier work this paper cites.
Data integration: After the teenage years. In Proceedings of the 36th ACM SIGMOD-SIGACT-SIGAI symposium on principles of database systems
Behzad Golshan, Alon Halevy, George Mihaila, and Wang-Chiew Tan. 2017 · 2017
Earlier work this paper cites.
A formal semantics of SQL queries, its validation, and applications
Paolo Guagliardo and Leonid Libkin. 2017 · 2017
Earlier work this paper cites.
Cross-sentence n-ary relation extraction with graph lstms
Nanyun Peng, Hoifung Poon, Chris Quirk, Kristina Toutanova, and Wen-tau Yih. 2017 · 2017
Earlier work this paper cites.
DeepDive: declarative knowledge base construction
Ce Zhang, Christopher Ré, Michael J. Cafarella, Jaeho Shin, Feiran Wang, and Sen Wu. 2017 · 2017
Earlier work this paper cites.
RHEEM: Enabling Cross-Platform Data Processing - May The Big Data Be With You! -
Divy Agrawal, Sanjay Chawla, Bertty Contreras-Rojas, Ahmed K. Elmagarmid, Yasser Idris, Zoi Kaoudi, Sebastian Kruse, Ji Lucas, Essam Mansour, Mourad Ouzzani, Paolo Papotti, Jorge-Arnulfo Quiané-Ruiz, Nan Tang, Saravanan Thirumuruganathan, and Anis Troudi. 2018 · 2018
Earlier work this paper cites.
SystemT: Declarative Text Understanding for Enterprise. In NAACL
Laura Chiticariu, Marina Danilevsky, Yunyao Li, Frederick Reiss, and Huaiyu Zhu. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task. In EMNLP
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Data Cleaning
Ihab F. Ilyas and Xu Chu. 2019 · 2019
Earlier work this paper cites.
Language Models as Knowledge Bases?. In EMNLP
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Auditing Data Provenance in Text-Generation Models. In KDD
Congzheng Song and Vitaly Shmatikov. 2019 · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners. In NeurIPS
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
TURL: Table Understanding through Representation Learning
Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu. 2020 · 2020
Earlier work this paper cites.
TaPas: Weakly Supervised Table Parsing via Pre-training. In ACL
Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Martin Eisenschlos. 2020 · 2020
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In NeurIPS
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Earlier work this paper cites.
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList. In ACL
Marco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Earlier work this paper cites.
How Much Knowledge Can You Pack Into the Parameters of a Language Model?. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2020
Earlier work this paper cites.
Break It Down: A Question Understanding Benchmark
Tomer Wolfson, Mor Geva, Ankit Gupta, Matt Gardner, Yoav Goldberg, Daniel Deutch, and Jonathan Berant. 2020 · 2020
Cited alongside, same era.
“Who said it, and Why?” Provenance for Natural Language Claims. In ACL
Yi Zhang, Zachary Ives, and Dan Roth. 2020 · 2020
Cited alongside, same era.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In FAccT
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Editing Factual Knowledge in Language Models. In EMNLP
Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021 · 2021
Cited alongside, same era.
Measuring and Improving Consistency in Pretrained Language Models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. 2022 · 2022
Later among the works it cites.
You are my type! Type embeddings for pre-trained language models. In EMNLP 2022, Conference on Empirical Methods in Natural Language Processing
Mohammed Saeed and Paolo Papotti. 2022 · 2022
Later among the works it cites.
Evaluating the Factual Consistency of Large Language Models Through Summarization
Derek Tam, Anisha Mascarenhas, Shiyue Zhang, Sarah Kwan, Mohit Bansal, and Colin Raffel. 2022 · 2022
Later among the works it cites.
Pythia: Unsupervised Generation of Ambiguous Textual Claims from Relational Data. In SIGMOD
Enzo Veltri, Donatello Santoro, Gilbert Badaro, Mohammed Saeed, and Paolo Papotti. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021 · 2021
Cited alongside, same era.
A Deep Dive into Deep Learning Approaches for Text-to-SQL Systems. In SIGMOD
George Katsogiannis-Meimarakis and Georgia Koutrika. 2021 · 2021
Cited alongside, same era.
Mind the Gap: Assessing Temporal Generalization in Neural Language Models. In NeurIPS
Angeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d’Autume, Tomás Kociský, Sebastian Ruder, Dani Yogatama, Kris Cao, Susannah Young, and Phil Blunsom. 2021 · 2021
Cited alongside, same era.
Capturing Semantics for Imputation with Pre-trained Language Models. In ICDE
Yinan Mei, Shaoxu Song, Chenguang Fang, Haifeng Yang, Jingyun Fang, and Jiang Long. 2021 · 2021
Cited alongside, same era.
From Natural Language Processing to Neural Databases
James Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri, Sebastian Riedel, and Alon Y. Levy. 2021 · 2021
Cited alongside, same era.
Factual Probing Is [MASK]: Learning vs. Learning to Recall. In North American Association for Computational Linguistics (NAACL)
Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021 · 2021
Cited alongside, same era.
The Seattle Report on Database Research
Daniel Abadi, Anastasia Ailamaki, David Andersen, Peter Bailis, Magdalena Balazinska, Philip A. Bernstein, Peter Boncz, Surajit Chaudhuri, Alvin Cheung, Anhai Doan, Luna Dong, Michael J. Franklin, Juliana Freire, Alon Halevy, Joseph M. Hellerstein, Stratos Idreos, Donald Kossmann, Tim Kraska, Sailesh Krishnamurthy, Volker Markl, Sergey Melnik, Tova Milo, C. Mohan, Thomas Neumann, Beng Chin Ooi, Fatma Ozcan, Jignesh Patel, Andrew Pavlo, Raluca Popa, Raghu Ramakrishnan, Christopher Re, Michael Stonebraker, and Dan Suciu. 2022 · 2022
Cited alongside, same era.
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc Le, and Denny Zhou. 2022 · 2022
Later among the works it cites.
Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation. In SIGMOD
Zixuan Zhao and Raul Castro Fernandez. 2022 · 2022
Later among the works it cites.
Introducing 100K Context Windows
Anthropic. 2023 · 2023
Closest in time.
Transformers for Tabular Data Representation: A Survey of Models and Applications
Gilbert Badaro, Mohammed Saeed, and Papotti Paolo. 2023 · 2023
Closest in time.
Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM
Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023 · 2023
Closest in time.
Entity Centric Neural Models for Natural Language Processing
Nicola De Cao. 2023 · 2023
Closest in time.
Textbooks Are All You Need
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li. 2023 · 2023
Closest in time.
Survey of Hallucination in Natural Language Generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Closest in time.
Pre-Train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023 · 2023
Closest in time.
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
Niels Mündler, Jingxuan He, Slobodan Jenko, and Martin Vechev. 2023 · 2023
Closest in time.
QATCH: Benchmarking Table Representation Learning Models on Your Data. In NeurIPS (Datasets and Benchmarks)
Simone Papicchio, Paolo Papotti, and Luca Cagliero. 2023 · 2023
Closest in time.
The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023 · 2023
Closest in time.
Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback
Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, and Jianfeng Gao. 2023 · 2023
Closest in time.
QA dataset explosion: A taxonomy of nlp resources for question answering and reading comprehension
Anna Rogers, Matt Gardner, and Isabelle Augenstein. 2023 · 2023
Closest in time.
Reflexion: Language Agents with Verbal Reinforcement Learning
Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023 · 2023
Closest in time.
BloombergGPT: A Large Language Model for Finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023 · 2023
Closest in time.