Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have gained significant attention in on-device applications due to their remarkable performance across real-world tasks.
Rethinking classification and localization for cascade r-cnn
Ang Li, Xue Yang, and Chongyang Zhang. 2019 · 1907
Earlier work this paper cites.
Rapid object detection using a boosted cascade of simple features
Paul Viola and Michael Jones. 2001 · 2001
Earlier work this paper cites.
A semisupervised cascade classification algorithm
Stamatis Karlos, Nikos Fazakis, Sotiris Kotsiantis, and Kyriakos Sgarbas. 2016 · 2016
Earlier work this paper cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Earlier work this paper cites.
Do prompt-based models really understand the meaning of their prompts?
Albert Webson and Ellie Pavlick. 2021 · 2021
Earlier work this paper cites.
Emailsum: Abstractive email thread summarization
Shiyue Zhang, Asli Celikyilmaz, Jianfeng Gao, and Mohit Bansal. 2021 · 2021
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Tryage: Real-time, intelligent routing of user prompts to large language model
Surya Narayanan Hari and Matt Thomson. 2023 · 2023
Earlier work this paper cites.
Cascaded machine learning model based dos attacks detection and classification in noc
Shengkai Hu, Haoyu Wang, and Basel Halak. 2023 · 2023
Earlier work this paper cites.
Preserving privacy through dememorization: An unlearning technique for mitigating memorization risks in language models
Aly Kassem, Omar Mahmoud, and Sherif Saad. 2023 · 2023
Earlier work this paper cites.
Orchestrallm: Efficient orchestration of language models for dialogue state tracking
Chia-Hsuan Lee, Hao Cheng, and Mari Ostendorf. 2023 · 2023
Earlier work this paper cites.
Reflection-tuning: Data recycling improves llm instruction-tuning
Ming Li, Lichang Chen, Jiuhai Chen, Shwai He, Heng Huang, Jiuxiang Gu, and Tianyi Zhou. 2023 · 2023
Earlier work this paper cites.
Automix: Automatically mixing language models
Aman Madaan, Pranjal Aggarwal, Ankit Anand, Srividya Pranavi Potharaju, Swaroop Mishra, Pei Zhou, Aditya Gupta, Dheeraj Rajagopal, Karthik Kappaganthu, Yiming Yang, et al. 2023 · 2023
Earlier work this paper cites.
Large language model routing with benchmark datasets
Tal Shnitzer, Anthony Ou, Mírian Silva, Kate Soule, Yuekai Sun, Justin Solomon, Neil Thompson, and Mikhail Yurochkin. 2023 · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Cited alongside, same era.
Tabi: An efficient multi-level inference system for large language models
Yiding Wang, Kai Chen, Haisheng Tan, and Kun Guo. 2023 · 2023
Cited alongside, same era.
Llmcad: Fast and scalable on-device large language model inference
Daliang Xu, Wangsong Yin, Xin Jin, Ying Zhang, Shiyun Wei, Mengwei Xu, and Xuanzhe Liu. 2023 · 2023
Cited alongside, same era.
Mars: A benchmark for multi-llm algorithmic routing system
Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Ranganath, Kurt Keutzer, and Shriyash Kaustubh Upadhyay. 2024a · 2024
Closest in time.
Preventing health data from leaking in a machine learning system: Implementing code analysis with llm and model privacy evaluation testing
Balder Janryd and Tim Johansson. 2024 · 2024
Closest in time.
When does confidence-based cascade deferral suffice?
Wittawat Jitkrittum, Neha Gupta, Aditya K Menon, Harikrishna Narasimhan, Ankit Rawat, and Sanjiv Kumar. 2024 · 2024
Closest in time.
Propile: Probing privacy leakage in large language models
Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh. 2024 · 2024
Closest in time.
Complying with hipaa and hitech
Susan Lincke. 2024 · 2024
Closest in time.
Llamoco: Instruction tuning of large language models for optimization code generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xue Yan, Yan Song, Xinyu Cui, Filippos Christianos, Haifeng Zhang, David Henry Mguni, and Jun Wang. 2023 · 2023
Cited alongside, same era.
Large language model cascades with mixture of thoughts representations for cost-efficient reasoning
Murong Yue, Jie Zhao, Min Zhang, Liang Du, and Ziyu Yao. 2023 · 2023
Cited alongside, same era.
Enhancing large language models’ utility for medical question-answering: A patient health question summarization approach
Nour Eddine Zekaoui, Siham Yousfi, Mounia Mikram, and Maryem Rhanoui. 2023 · 2023
Cited alongside, same era.
Towards llm-based fact verification on news claims with a hierarchical step-by-step prompting method
Xuan Zhang and Wei Gao. 2023 · 2023
Cited alongside, same era.
Offline multi-action policy learning: Generalization and optimization
Zhengyuan Zhou, Susan Athey, and Stefan Wager. 2023 · 2023
Cited alongside, same era.
Vidur: A large-scale simulation framework for llm inference
Amey Agrawal, Nitin Kedia, Jayashree Mohan, Ashish Panwar, Nipun Kwatra, Bhargav Gulavani, Ramachandran Ramjee, and Alexey Tumanov. 2024 · 2024
Cited alongside, same era.
Boyuan Chen, Mingzhi Zhu, Brendan Dolan-Gavitt, Muhammad Shafique, and Siddharth Garg. 2024 · 2024
Cited alongside, same era.
Security and privacy challenges of large language models: A survey
Badhan Chandra Das, M Hadi Amini, and Yanzhao Wu. 2024 · 2024
Cited alongside, same era.
Zeyuan Ma, Hongshu Guo, Jiacheng Chen, Guojun Peng, Zhiguang Cao, Yining Ma, and Yue-Jiao Gong. 2024 · 2024
Closest in time.
Pocketllm: Enabling on-device fine-tuning for personalized llms
Dan Peng, Zhihui Fu, and Jun Wang. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. 2024 · 2024
Closest in time.
Fly-swat or cannon? cost-effective language model choice via meta-modeling
Marija Šakota, Maxime Peyrard, and Robert West. 2024 · 2024
Closest in time.
Towards optimizing the costs of llm usage
Shivanshu Shekhar, Tanishq Dubey, Koyel Mukherjee, Apoorv Saxena, Atharv Tyagi, and Nishanth Kotla. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. 2024 · 2024
Closest in time.
Cascade-aware training of language models
Congchao Wang, Sean Augenstein, Keith Rush, Wittawat Jitkrittum, Harikrishna Narasimhan, Ankit Singh Rawat, Aditya Krishna Menon, and Alec Go. 2024 · 2024
Closest in time.
Rethinking chain-of-thought from the perspective of self-training
Zongqian Wu, Baoduo Xu, Ruochen Cui, Mengmeng Zhan, Xiaofeng Zhu, and Lei Feng. 2024 · 2024
Closest in time.
On-device language models: A comprehensive review
Jiajun Xu, Zhiyuan Li, Wei Chen, Qun Wang, Xin Gao, Qi Cai, and Ziyuan Ling. 2024 · 2024
Closest in time.
Wip: An on-device llm-based approach to query privacy protection
Yizhen Yuan, Rui Kong, Yuanchun Li, and Yunxin Liu. 2024 · 2024
Closest in time.
LLM-based medical assistant personalization with short- and long-term memory coordination
Kai Zhang, Yangyang Kang, Fubang Zhao, and Xiaozhong Liu. 2024a · 2024
Closest in time.
A novel cascade instruction tuning method for biomedical ner
Jin Zhao, Chao Liu, Jiaqing Liang, Zhixu Li, and Yanghua Xiao. 2024 · 2024
Closest in time.
Towards an on-device agent for text rewriting
Yun Zhu, Yinxiao Liu, Felix Stahlberg, Shankar Kumar, Yu-Hui Chen, Liangchen Luo, Lei Shu, Renjie Liu, Jindong Chen, and Lei Meng. 2024 · 2024
Closest in time.