Fetching the paper…
Reading the bibliography…
A long-standing goal of data management systems has been to build systems which can compute quantitative insights over large corpora of unstructured data in a cost-effective manner.
Potter’s wheel: An interactive data cleaning system
Vijayshankar Raman and Joseph M. Hellerstein · 2001
Earlier work this paper cites.
Applying model management to classical meta data problems
Philip A Bernstein · 2003
Earlier work this paper cites.
The design of an acquisitional query processor for sensor networks
Samuel Madden, Michael J. Franklin, Joseph M. Hellerstein, and Wei Hong · 2003
Earlier work this paper cites.
Crowddb: Query processing with the VLDB crowd
Amber Feng, Michael J. Franklin, Donald Kossmann, Tim Kraska, Samuel Madden, Sukriti Ramesh, Andrew Wang, and Reynold Xin · 2011
Earlier work this paper cites.
Crowdsourced databases: Query processing with people
Adam Marcus, Eugene Wu, Samuel Madden, and Robert C. Miller · 2011
Earlier work this paper cites.
Counting with the crowd
Adam Marcus, David R. Karger, Samuel Madden, Rob Miller, and Sewoong Oh · 2012
Earlier work this paper cites.
Deco: A system for declarative crowdsourcing
Hyunjung Park, Richard Pang, Aditya G. Parameswaran, Hector Garcia-Molina, Neoklis Polyzotis, and Jennifer Widom · 2012
Earlier work this paper cites.
Deep convolutional network cascade for facial point detection
Yi Sun, Xiaogang Wang, and Xiaoou Tang · 2013
Earlier work this paper cites.
Learning complexity-aware cascades for deep pedestrian detection
Zhaowei Cai, Mohammad J. Saberian, and Nuno Vasconcelos · 2015
Earlier work this paper cites.
Enron email dataset, May 2015
William W. Cohen · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Noscope: optimizing neural network queries over video at scale
Daniel Kang, John Emmons, Firas Abuzaid, Peter Bailis, and Matei Zaharia · 2017
Earlier work this paper cites.
Snorkel: rapid training data creation with weak supervision
Alexander Ratner, Stephen H. Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher Ré · 2017
Earlier work this paper cites.
Physical representation-based predicate optimization for a visual analytics database
Michael R Anderson, Michael Cafarella, German Ros, and Thomas F Wenisch · 2019
Earlier work this paper cites.
Proteogenomic characterization of ovarian hgsc implicates mitotic kinases, replication stress in observed chromosomal instability
Jason E. McDermott, Osama A. Arshad, Vladislav A. Petyuk, Yi Fu, Marina A. Gritsenko, Therese R. Clauss, Ronald J. Moore, Athena A. Schepmoes, Rui Zhao, Matthew E. Monroe, Michael Schnaubelt, Chia-Feng Tsai, Samuel H. Payne, Chen Huang, Liang-Bo Wang, Steven Foltz, Matthew Wyczalkowski, Yige Wu, Ehwang Song, Molly A. Brewer, Mathangi Thiagarajan, Christopher R. Kinsinger, Ana I. Robles, Emily S. Boja, Henry Rodriguez, Daniel W. Chan, Bing Zhang, Zhen Zhang, Li Ding, Richard D. Smith, Tao Liu, and Karin D. Rodland · 2020
Earlier work this paper cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J. Maybank, and Dacheng Tao · 2021
Earlier work this paper cites.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals · 2022
Earlier work this paper cites.
LlamaIndex, 11 2022
Jerry Liu · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Language models enable simple systems for generating structured views of heterogeneous data lakes
Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel Trummer, and Christopher Ré · 2023
Cited alongside, same era.
Accelerating large language model decoding with speculative sampling
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper · 2023
Cited alongside, same era.
Frugalgpt: How to use large language models while reducing cost and improving performance
Lingjiao Chen, Matei Zaharia, and James Zou · 2023
Cited alongside, same era.
Seed: Simple, efficient, and effective data management via large language models
Zui CHen, Lei Cao, Sam Madden, Ju Fan, Nan Tang, Zihui Gu, Zeyuan Shang, Chunwei Liu, Michael Cafarella, and Tim Kraska · 2023
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang · 2023
Later among the works it cites.
Skypilot: An intercloud broker for sky computing
Zongheng Yang, Zhanghao Wu, Michael Luo, Wei-Lin Chiang, Romil Bhardwaj, Woosuk Kwon, Siyuan Zhuang, Frank Sifei Luan, Gautam Mittal, Scott Shenker, and Ion Stoica · 2023
Later among the works it cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2023
Later among the works it cites.
Efficiently programming large language models using sglang
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al · 2023
Later among the works it cites.
https://modal.com , 2024
Modal.com api · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Minillm: Knowledge distillation of large language models
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang · 2023
Cited alongside, same era.
Guidance AI: Open source project for AI development
Guidance AI Contributors · 2023
Cited alongside, same era.
Metagpt: Meta programming for multi-agent collaborative framework
Sirui Hong, Xiawu Zheng, Jonathan Chen, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, et al · 2023
Cited alongside, same era.
Dspy: Compiling declarative language model calls into self-improving pipelines
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T Joshi, Hanna Moazam, et al · 2023
Cited alongside, same era.
Validating large language models with relm
Michael Kuchnik, Virginia Smith, and George Amvrosiadis · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Cited alongside, same era.
Proteogenomic data and resources for pan-cancer analysis
Yize Li, Yongchao Dou, Felipe Da Veiga Leprevost, Yifat Geffen, Anna P Calinawan, François Aguet, Yo Akiyama, Shankara Anand, Chet Birger, Song Cao, et al · 2023
Cited alongside, same era.
Closest in time.
https://ollama.com , 2024
Ollama · 2024
Closest in time.
https://together.ai , 2024
Together.ai · 2024
Closest in time.
CascadeServe: Unlocking model cascades for inference serving
Anonymous authors · 2024
Closest in time.
Langchain: Open source framework for building language models
LangChain Contributors · 2024
Closest in time.
Towards accurate and efficient document analytics with large language models
Yiming Lin, Madelon Hulsebos, Ruiying Ma, Shreya Shankar, Sepanta Zeigham, Aditya G Parameswaran, and Eugene Wu · 2024
Closest in time.
Optimizing llm queries in relational workloads
Shu Liu, Asim Biswal, Audrey Cheng, Xiangxi Mo, Shiyi Cao, Joseph E Gonzalez, Ion Stoica, and Matei Zaharia · 2024
Closest in time.
Openai api
OpenAI · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2024
Closest in time.
Cascade inference: Memory bandwidth efficient shared prefix batch decoding
Zihao Ye, Ruihang Lai, Bo-Ru Lu, Chien-Yu Lin, Size Zheng, Lequn Chen, Tianqi Chen, and Luis Ceze · 2024
Closest in time.
The shift from models to compound ai systems
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi · 2024
Closest in time.
End-to-end beam retrieval for multi-hop question answering
Jiahao Zhang, Haiyang Zhang, Dongmei Zhang, Yong Liu, and Shen Huang · 2024
Closest in time.
Prepacking: A simple method for fast prefilling and increased throughput in large language models
Siyan Zhao, Daniel Israel, Guy Van den Broeck, and Aditya Grover · 2024
Closest in time.