Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are often praised for exhibiting near-human performance on a wide range of tasks and valued for their ability to hold a general conversation.
Brain size and ecology in small mammals
Georgina M Mace, Paul H Harvey, and Timothy H Clutton-Brock · 1981
Earlier work this paper cites.
Microfossils of the early archean apex chert: New evidence of the antiquity of life
J William Schopf · 1993
Earlier work this paper cites.
Before the Beginning: Our Universe and Others
Martin J Rees · 1997
Earlier work this paper cites.
Coarse-scale population structure of pathogenic armillaria species in a mixed-conifer forest in the blue mountains of northeast oregon
Barbara A Ferguson, Timothy A Dreisbach, Catherine G Parks, Gregory M Filip, and Craig L Schmitt · 2003
Earlier work this paper cites.
A comparative study on unsupervised feature selection methods for text clustering
Luying Liu, Jianchu Kang, Jing Yu, and Zhongliang Wang · 2005
Earlier work this paper cites.
Adversarial stylometry: Circumventing authorship recognition to preserve privacy and anonymity
Michael Brennan, Sadia Afroz, and Rachel Greenstadt · 2012
Earlier work this paper cites.
Planck 2018 results. vi. cosmological parameters
Planck Collaboration et al · 2020
Earlier work this paper cites.
Unsupervised fine-tuning for text clustering
Shaohan Huang, Furu Wei, Lei Cui, Xingxing Zhang, and Ming Zhou · 2020
Earlier work this paper cites.
Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models. arxiv 2021
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Bogdan Damoc, Aidan Clark, Jan Kramár, et al · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2022
Earlier work this paper cites.
Least-to-most prompting enables complex reasoning in small language models
Xuezhi Zhou, Nathanael Schärli, Yujie Hou, Jason Wei, Denny Zhou, Quoc V. Le, and Douwe Kiela · 2022
Earlier work this paper cites.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Earlier work this paper cites.
Parameter-efficient fine-tuning of large-scale pre-trained language models
Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al · 2023
Earlier work this paper cites.
Phi-2: The surprising power of small language models, 2023
Mojan Javaheripi and Sébastien Bubeck · 2023
Earlier work this paper cites.
Artificial intelligence and democracy: A conceptual framework
Andreas Jungherr · 2023
Earlier work this paper cites.
Matformer: Nested transformer for elastic inference
Sneha Kudugunta, Aditya Kusupati, Tim Dettmers, Kaifeng Chen, Inderjit Dhillon, Yulia Tsvetkov, Hannaneh Hajishirzi, Sham Kakade, Ali Farhadi, Prateek Jain, et al · 2023
Earlier work this paper cites.
Deja vu: Contextual sparsity for efficient llms at inference time
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al · 2023
Earlier work this paper cites.
A comprehensive overview of large language models
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian · 2023
Earlier work this paper cites.
A certified de-identification system for all clinical text documents for information extraction at scale
Lakshmi Radhakrishnan, Gundolf Schenk, Kathleen Muenzen, Boris Oskotsky, Habibeh Ashouri Choshali, Thomas Plunkett, Sharat Israni, and Atul J Butte · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Earlier work this paper cites.
Instruction-following evaluation for large language models
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou · 2023
Earlier work this paper cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al · 2024
Earlier work this paper cites.
Delift: Data efficient language model instruction fine tuning
Ishika Agarwal, Krishnateja Killamsetty, Lucian Popa, and Marina Danilevksy · 2024
Earlier work this paper cites.
Tiny transformers excel at sentence compression
Peter Belcak and Roger Wattenhofer · 2024
Earlier work this paper cites.
Flextron: Many-in-one flexible large language model
Ruisi Cai, Saurav Muralidharan, Greg Heinrich, Hongxu Yin, Zhangyang Wang, Jan Kautz, and Pavlo Molchanov · 2024
Earlier work this paper cites.
Hymba: A hybrid-head architecture for small language models
Xin Dong, Yonggan Fu, Shizhe Diao, Wonmin Byeon, Zijia Chen, Ameya Sunil Mahabaleshwarkar, Shih-Yang Liu, Matthijs Van Keirsbilck, Min-Hung Chen, Yoshi Suhara, et al · 2024
Cited alongside, same era.
Amoeballm: Constructing any-shape large language models for efficient and instant deployment
Yonggan Fu, Zhongzhi Yu, Junwei Li, Jiayi Qian, Yongan Zhang, Xiangchi Yuan, Dachuan Shi, Roman Yakunin, and Yingyan Celine Lin · 2024
Cited alongside, same era.
Dora: Weight-decomposed low-rank adaptation
Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, and Min-Hung Chen · 2024
Cited alongside, same era.
Nunzio Lore, Sepehr Ilami, and Babak Heydari · 2024
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.
Climb: Clustering-based iterative data mixture bootstrapping for language model pre-training
Shizhe Diao, Yu Yang, Yonggan Fu, Xin Dong, Dan Su, Markus Kliegl, Zijia Chen, Peter Belcak, Yoshi Suhara, Hongxu Yin, et al · 2025
Closest in time.
Introducing nvidia dynamo, a low-latency distributed inference framework for scaling reasoning ai models, March 2025
Amr Elmeleegy et al · 2025
Closest in time.
Llms vs. slms: Balancing comprehensiveness and smart resource-saving, April 2025
Henry Evans · 2025
Closest in time.
Text compression for efficient language generation
David Gu, Peter Belcak, and Roger Wattenhofer · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhenyan Lu, Xiang Li, Dongqi Cai, Rongjie Yi, Fangming Liu, Xiwen Zhang, Nicholas D Lane, and Mengwei Xu · 2024
Cited alongside, same era.
The landscape of emerging ai agent architectures for reasoning, planning, and tool calling: A survey
Tula Masterman, Sandi Besen, Mason Sawtell, and Alex Chao · 2024
Cited alongside, same era.
Chatrtx, 2024
NVIDIA · 2024
Cited alongside, same era.
tinybenchmarks: evaluating llms with fewer examples
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin · 2024
Cited alongside, same era.
An open source python library for anonymizing sensitive data
Judith Sáinz-Pardo Díaz and Álvaro López García · 2024
Cited alongside, same era.
Powerinfer: Fast large language model serving with a consumer-grade gpu
Yixin Song, Zeyu Mi, Haotong Xie, and Haibo Chen · 2024
Cited alongside, same era.
Fali Wang, Zhiwei Zhang, Xianren Zhang, Zongyu Wu, Tzuhao Mo, Qiuhao Lu, Wanjing Wang, Rui Li, Junjie Xu, Xianfeng Tang, et al · 2024
Cited alongside, same era.
Powerinfer-2: Fast large language model inference on a smartphone
Zhenliang Xue, Yixin Song, Zeyu Mi, Xinrui Zheng, Yubin Xia, and Haibo Chen · 2024
Cited alongside, same era.
Large language models vs. small language models, March 2024
Harrison Clarke · 2025
Closest in time.
How small language models can outperform llms, March 2025
Invisible Technologies · 2025
Closest in time.
Understanding the total cost of inferencing large language models
Aviv Kaufmann · 2025
Closest in time.
From large to small: The rise of small language models (slms) in text analytics
Akshi Kumar · 2025
Closest in time.
Autonomous generative ai agents: Under development
Jeff Loucks, Gillian Crossan, Baris Sarer, China Widener, and Ariane Bucaille · 2025
Closest in time.
Large language model agent: A survey on methodology, applications and challenges
Junyu Luo, Weizhi Zhang, Ye Yuan, Yusheng Zhao, Junwei Yang, Yiyang Gu, Bohan Wu, Binqi Chen, Ziyue Qiao, Qingqing Long, et al · 2025
Closest in time.
A closer look at dynamo, nvidia’s ’operating system’ for ai inference, March 2025
Tobias Mann · 2025
Closest in time.
Global agentic ai market size, share analysis by product type, agent role, agent system, end user, region and companies – industry segment outlook, market assessment, competition scenario, trends and forecast 2025–2034, March 2025
Market.us · 2025
Closest in time.
How much energy do llms consume? unveiling the power behind ai, July 2024
Sourabh Mehta · 2025
Closest in time.
Model cards and prompt formats: Llama 3.3, April 2025
Meta Platforms, Inc · 2025
Closest in time.
Understanding ai agents & data security, 2025
Metomic · 2025
Closest in time.
Agentic ai needs a systems theory
Erik Miehling, Karthikeyan Natesan Ramamurthy, Kush R Varshney, Matthew Riemer, Djallel Bouneffouf, John T Richards, Amit Dhurandhar, Elizabeth M Daly, Michael Hind, Prasanna Sattigeri, et al · 2025
Closest in time.
Genai revenue growth and profitability, April 2025
Morgan Stanley · 2025
Closest in time.
Nvidia dynamo: A datacenter scale distributed inference serving framework
NVIDIA · 2025
Closest in time.
Cloud llm cost model: Breakdown for mid-market businesses, 2024
Tanya Seda · 2025
Closest in time.
Explore ai models: Key differences between small language models and large language models, November 2024
Olivia Shone · 2025
Closest in time.
Small language models (slms) can still pack a punch: A survey
Shreyas Subramanian, Vikram Elango, and Mecit Gungor · 2025
Closest in time.
Small language models vs. large language models, 2025
Synergy Technical · 2025
Closest in time.
Trustworthy and secure ai: How small language models strengthen data security
Brian G. Thamm · 2025
Closest in time.
Build secure ai agents, 2025
WorkOS · 2025
Closest in time.
Like human brains, large language models reason about diverse data in a general way
Adam Zewe · 2025
Closest in time.
A deep dive on ai inference startups, 2024
Kevin Zhang · 2025
Closest in time.
Introducing nvidia dynamo, a low-latency distributed inference framework for scaling reasoning ai models, March 2025
David Zier and Harry Kim · 2025
Closest in time.