Fetching the paper…
Reading the bibliography…
State-of-the-art sequential reasoning in Large Language Models (LLMs) has expanded the capabilities of Copilots beyond conversational tasks to complex function calling, managing thousands of API calls.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021 · 2005
Earlier work this paper cites.
Neuralpower: Predict and deploy energy-efficient convolutional neural networks. In Asian Conference on Machine Learning . PMLR, 622–637
Ermao Cai, Da-Cheng Juan, Dimitrios Stamoulis, and Diana Marculescu. 2017 · 2017
Earlier work this paper cites.
Hardware-aware machine learning: Modeling and optimization. In 2018 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) . IEEE, 1–8
Diana Marculescu, Dimitrios Stamoulis, and Ermao Cai. 2018 · 2018
Earlier work this paper cites.
A Fast Post-Training Pruning Framework for Transformers
Woosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami. 2022 · 2022
Earlier work this paper cites.
Adapting Language Models to Compress Contexts
Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023 · 2023
Earlier work this paper cites.
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. 2023 · 2023
Earlier work this paper cites.
A Survey on In-context Learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui. 2023 · 2023
Earlier work this paper cites.
Extending Context Window of Large Language Models via Semantic Compression
Weizhi Fei, Xueyan Niu, Pingyi Zhou, Lu Hou, Bo Bai, Lei Deng, and Wei Han. 2023 · 2023
Earlier work this paper cites.
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2023 · 2023
Earlier work this paper cites.
LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023 · 2023
Earlier work this paper cites.
OmniBoost: Boosting Throughput of Heterogeneous Embedded Devices under Multi-DNN Workload. In 2023 60th ACM/IEEE Design Automation Conference (DAC) . IEEE, 1–6
Andreas Karatzas and Iraklis Anagnostopoulos. 2023 · 2023
Earlier work this paper cites.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
Compressing Context to Enhance Inference Efficiency of Large Language Models
Yucheng Li, Bo Dong, Chenghua Lin, and Frank Guerin. 2023 · 2023
Earlier work this paper cites.
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2023 · 2023
Earlier work this paper cites.
NexusRaven: a commercially-permissive Language Model for function calling. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following
Venkat Krishna Srinivasan, Zhen Dong, Banghua Zhu, Brian Yu, Hanzi Mao, Damon Mosk-Aoyama, Kurt Keutzer, Jiantao Jiao, and Jian Zhang. 2023 · 2023
Cited alongside, same era.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023 · 2023
Cited alongside, same era.
RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentation
Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2023 · 2023
Cited alongside, same era.
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang. 2023 · 2023
Cited alongside, same era.
Retrieval Augmented Generation (RAG)
LangChain Docs. 2024 · 2024
Closest in time.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024 · 2024
Closest in time.
APAR: LLMs Can Do Auto-Parallel Auto-Regressive Decoding
Mingdao Liu, Aohan Zeng, Bowen Wang, Peng Zhang, Jie Tang, and Yuxiao Dong. 2024 · 2024
Closest in time.
TOFU: A Task of Fictitious Unlearning for LLMs
Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C. Lipton, and J. Zico Kolter. 2024 · 2024
Closest in time.
OpenEQA: Embodied Question Answering in the Era of Foundation Models. In Conference on Computer Vision and Pattern Recognition (CVPR)
Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma, Vincent Berges, Shiqi Zhang, Pulkit Agrawal, Yonatan Bisk, Dhruv Batra, Mrinal Kalakrishnan, Franziska Meier, Chris Paxton, Sasha Sax, and Aravind Rajeswaran. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2023 · 2023
Cited alongside, same era.
Efficient Prompting via Dynamic In-Context Learning
Wangchunshu Zhou, Yuchen Eleanor Jiang, Ryan Cotterell, and Mrinmaya Sachan. 2023 · 2023
Cited alongside, same era.
ToolQA: A Dataset for LLM Question Answering with External Tools
Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang. 2023 · 2023
Cited alongside, same era.
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Abdin M. et al. 2024 · 2024
Cited alongside, same era.
Introducing Meta Llama 3: The most capable openly available LLM to date
Meta AI. 2024 · 2024
Cited alongside, same era.
Hardware-Aware DNN Compression via Diverse Pruning and Mixed-Precision Quantization
Konstantinos Balaskas, Andreas Karatzas, Christos Sad, Kostas Siozios, Iraklis Anagnostopoulos, Georgios Zervakis, et al · 2024
Cited alongside, same era.
Text Generation Inference: A Rust, Python, and gRPC server for text generation inference
Hugging Face. 2024 · 2024
Cited alongside, same era.
GeckOpt: LLM System Efficiency via Intent-Based Tool Selection. In GLSVLSI 2024
Michael Fore, Simranjit Singh, and Dimitrios Stamoulis. 2024 · 2024
Cited alongside, same era.
Closest in time.
Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation
Xuefei Ning, Zinan Lin, Zixuan Zhou, Zifu Wang, Huazhong Yang, and Yu Wang. 2024 · 2024
Closest in time.
TensorRT-LLM: A TensorRT Toolbox for Optimized Large Language Model Inference
NVIDIA. 2024 · 2024
Closest in time.
Function Calling
OpenAI API Docs. 2024 · 2024
Closest in time.
Parallel Function Calling
OpenAI Developer community. 2024 · 2024
Closest in time.
Evaluating Tool-Augmented Agents in Remote Sensing Platforms. In ICLR 2024 Workshop: 2nd Machine Learning for Remote Sensing Workshop
Simranjit Singh, Michael Fore, and Dimitrios Stamoulis. 2024a · 2024
Closest in time.
GeoLLM-Engine: A Realistic Environment for Building Geospatial Copilots. In CVPR 2024 Workshop EARTHVISION 2024
Simranjit Singh, Michael Fore, and Dimitrios Stamoulis. 2024b · 2024
Closest in time.
Berkeley Function Calling Leaderboard
Fanjia Yan, Huanzhi Mao, Charlie Cheng-Jie Ji, Tianjun Zhang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. 2024 · 2024
Closest in time.
LLM Inference Unveiled: Survey and Roofline Model Insights
Zhihang Yuan, Yuzhang Shang, Yang Zhou, Zhen Dong, Zhe Zhou, Chenhao Xue, Bingzhe Wu, Zhikai Li, Qingyi Gu, Yong Jae Lee, Yan Yan, Beidi Chen, Guangyu Sun, and Kurt Keutzer. 2024 · 2024
Closest in time.