Fetching the paper…
Reading the bibliography…
Transformer-based Large Language Models (LLMs) have been applied in diverse areas such as knowledge bases, human interfaces, and dynamic agents, and marking a stride towards achieving Artificial General Intelligence (AGI).
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020 · 1909
Earlier work this paper cites.
One-level storage system
Tom Kilburn, David BG Edwards, Michael J Lanigan, and Frank H Sumner. 1962 · 1962
Earlier work this paper cites.
The Logical Form of Action Sentences
Donald Davidson. 1967 · 1967
Earlier work this paper cites.
Virtual memory
Peter J Denning. 1970 · 1970
Earlier work this paper cites.
Adaptive multivariate ridge regression
Philip J Brown and James V Zidek. 1980 · 1980
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. 1991 · 1991
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998 · 1998
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht. 2007 · 2007
Earlier work this paper cites.
Meteor, m-bleu and m-ter: Evaluation metrics for high-correlation with human rankings of machine translation output. In Proceedings of the Third Workshop on Statistical Machine Translation . 115–118
Abhaya Agarwal and Alon Lavie. 2008 · 2008
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky. 2009 · 2009
Earlier work this paper cites.
Efficient Transformers: A Survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2022 · 2009
Earlier work this paper cites.
Tiling for performance tuning on different models of GPUs. In 2009 Second International Symposium on Information Science and Engineering . IEEE, 500–504
Chang Xu, Steven R Kirk, and Samantha Jenkins. 2009 · 2009
Earlier work this paper cites.
Recurrent neural network based language model.. In Interspeech , Vol. 2. Makuhari, 1045–1048
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Naive Bayes
Geoffrey I Webb, Eamonn Keogh, and Risto Miikkulainen. 2010 · 2010
Earlier work this paper cites.
Robust principal component analysis?
Emmanuel J Candès, Xiaodong Li, Yi Ma, and John Wright. 2011 · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
Context dependent recurrent neural network language model. In 2012 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 234–239
Tomas Mikolov and Geoffrey Zweig. 2012 · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter. 2015 · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015 · 2015
Earlier work this paper cites.
Multi-scale context aggregation by dilated convolutions
Fisher Yu and Vladlen Koltun. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh. 2016 · 2016
Earlier work this paper cites.
Hierarchical attention networks for document classification. In Proceedings of the 2016 conference of the North American chapter of the association for computational linguistics: human language technologies
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016 · 2016
Earlier work this paper cites.
The unreasonable effectiveness of structured random orthogonal embeddings
Krzysztof M Choromanski, Mark Rowland, and Adrian Weller. 2017 · 2017
Earlier work this paper cites.
The NarrativeQA Reading Comprehension Challenge
Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2017 · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Graph attention networks
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, Yoshua Bengio, et al · 2017
Earlier work this paper cites.
A discourse-aware attention model for abstractive summarization of long documents
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018 · 2018
Earlier work this paper cites.
Entity mention aware document representation
Hongliang Dai, Siliang Tang, Fei Wu, and Yueting Zhuang. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Rouge 2.0: Updated and improved measures for evaluation of summarization tasks
Kavita Ganesan. 2018 · 2018
Earlier work this paper cites.
Pipedream: Fast and efficient pipeline parallel dnn training
Aaron Harlap, Deepak Narayanan, Amar Phanishayee, Vivek Seshadri, Nikhil Devanur, Greg Ganger, and Phil Gibbons. 2018 · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler. 2018 · 2018
Earlier work this paper cites.
Dissecting the NVIDIA volta GPU architecture via microbenchmarking
Zhe Jia, Marco Maggioni, Benjamin Staiger, and Daniele P Scarpazza. 2018 · 2018
Earlier work this paper cites.
Sharp nearby, fuzzy far away: How neural language models use context
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Earlier work this paper cites.
Darts: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang. 2018 · 2018
Earlier work this paper cites.
Document-level neural machine translation with hierarchical attention networks
Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018 · 2018
Earlier work this paper cites.
Online normalizer calculation for softmax
Maxim Milakov and Natalia Gimelshein. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Earlier work this paper cites.
Bi-directional block self-attention for fast and memory-efficient sequence modeling
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, and Chengqi Zhang. 2018 · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Earlier work this paper cites.
Deep learning for sentiment analysis: A survey
Lei Zhang, Shuai Wang, and Bing Liu. 2018 · 2018
Earlier work this paper cites.
Hierarchical attentional hybrid neural networks for document classification. In International Conference on Artificial Neural Networks . Springer, 396–402
Jader Abreu, Luis Fred, David Macêdo, and Cleber Zanchettin. 2019 · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 2019
Earlier work this paper cites.
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Yunfan Shao, Xiangyang Xue, and Zheng Zhang. 2019 · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Earlier work this paper cites.
Mnnfast: A fast and scalable system architecture for memory-augmented neural networks. In Proceedings of the 46th International Symposium on Computer Architecture . 250–263
Hanhwi Jang, Joonsung Kim, Jae-Eon Jo, Jaewon Lee, and Jangwoo Kim. 2019 · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 2019
Earlier work this paper cites.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang, and Xifeng Yan. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Earlier work this paper cites.
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Scott Wen-tau Yih, Sinong Wang, and Jie Tang. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap. 2019 · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Noam Shazeer. 2019 · 2019
Earlier work this paper cites.
Low-memory neural network training: A technical report
Nimit S Sohoni, Christopher R Aberger, Megan Leszczynski, Jian Zhang, and Christopher Ré. 2019 · 2019
Earlier work this paper cites.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. 2019 · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages . 10–19
Philippe Tillet, Hsiang-Tsung Kung, and David Cox. 2019 · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Earlier work this paper cites.
Bp-transformer: Modelling long-range context via binary partitioning
Zihao Ye, Qipeng Guo, Quan Gan, Xipeng Qiu, and Zheng Zhang. 2019 · 2019
Earlier work this paper cites.
A review of recurrent neural networks: LSTM cells and network architectures
Yong Yu, Xiaosheng Si, Changhua Hu, and Jianxun Zhang. 2019 · 2019
Earlier work this paper cites.
Q8bert: Quantized 8bit bert. In 2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition (EMC2-NIPS) . IEEE, 36–39
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019 · 2019
Earlier work this paper cites.
Xingxing Zhang, Furu Wei, and Ming Zhou. 2019 · 2019
Earlier work this paper cites.
ETC: Encoding long and structured inputs in transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020 · 2020
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020 · 2020
Earlier work this paper cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Earlier work this paper cites.
Funnel-transformer: Filtering out sequential redundancy for efficient language processing
Zihang Dai, Guokun Lai, Yiming Yang, and Quoc Le. 2020 · 2020
Earlier work this paper cites.
Cogltx: Applying bert to long texts
Ming Ding, Chang Zhou, Hongxia Yang, and Jie Tang. 2020b · 2020
Earlier work this paper cites.
ERNIE-Doc: A retrospective long-document modeling transformer
Siyu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2020a · 2020
Earlier work this paper cites.
Addressing some limitations of transformers with feedback memory
Angela Fan, Thibaut Lavril, Edouard Grave, Armand Joulin, and Sainbayar Sukhbaatar. 2020 · 2020
Earlier work this paper cites.
A divide-and-conquer approach to the summarization of long documents
Alexios Gidiotis and Grigorios Tsoumakas. 2020 · 2020
Earlier work this paper cites.
REALM: Retrieval augmented language model pre-training. In International conference on machine learning . PMLR, 3929–3938
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020 · 2020
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering
Gautier Izacard and Edouard Grave. 2020 · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention. In International conference on machine learning . PMLR, 5156–5165
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Pytorch distributed: Experiences on accelerating data parallel training
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, et al · 2020
Cited alongside, same era.
Sparse and continuous attention mechanisms
André Martins, António Farinhas, Marcos Treviso, Vlad Niculae, Pedro Aguiar, and Mario Figueiredo. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 1–16
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Cited alongside, same era.
Sparse sinkhorn attention. In International Conference on Machine Learning . PMLR, 9438–9447
L-Eval: Instituting Standardized Evaluation for Long Context Language Models
Chenxin An, Shansan Gong, Ming Zhong, Mukai Li, Jun Zhang, Lingpeng Kong, and Xipeng Qiu. 2023 · 2023
Closest in time.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Closest in time.
Model Card and Evaluations for Claude Models
Anthropic. 2023 · 2023
Closest in time.
Half-Quadratic Quantization of Large Machine Learning Models
Hicham Badri and Appu Shaji. 2023 · 2023
Closest in time.
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan. 2020 · 2020
Cited alongside, same era.
DeepSpeed: Extreme-scale model training for everyone
DeepSpeed Team and Rangan Majumder. 2020 · 2020
Cited alongside, same era.
How to generate text: using different decoding methods for language generation with Transformers
Patrick von Platen. 2020 · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. 2020a · 2020
Cited alongside, same era.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020b · 2020
Cited alongside, same era.
Memformer: A memory-augmented transformer for sequence modeling
Qingyang Wu, Zhenzhong Lan, Kun Qian, Jing Gu, Alborz Geramifard, and Zhou Yu. 2020a · 2020
Cited alongside, same era.
Lite transformer with long-short range attention
Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han. 2020b · 2020
Cited alongside, same era.
O (n) connections are expressive enough: Universal approximability of sparse transformers
Chulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar. 2020 · 2020
Cited alongside, same era.
Closest in time.
Alternating Updates for Efficient Transformers
Cenk Baykal, Dylan Cutler, Nishanth Dikkala, Nikhil Ghosh, Rina Panigrahy, and Xin Wang. 2023 · 2023
Closest in time.
Unlimiformer: Long-range transformers with unlimited length input
Amanda Bertsch, Uri Alon, Graham Neubig, and Matthew R Gormley. 2023 · 2023
Closest in time.
Add NTK-Aware interpolation "by parts" correction, 2023
bloc97. 2023a · 2023
Closest in time.
Sparks of Artificial General Intelligence: Early experiments with GPT-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023 · 2023
Closest in time.
Scaling Transformer to 1M tokens and beyond with RMT
Aydar Bulatov, Yuri Kuratov, and Mikhail S Burtsev. 2023 · 2023
Closest in time.
Langchain-Chatchat: A LLM application aims to implement knowledge and search engine based QA based on Langchain and open-source or remote LLM API
Chatchat-Space. 2023 · 2023
Closest in time.
CLEX: Continuous Length Extrapolation for Large Language Models
Guanzheng Chen, Xin Li, Zaiqiao Meng, Shangsong Liang, and Lidong Bing. 2023b · 2023
Closest in time.
Extending context window of large language models via positional interpolation
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023e · 2023
Closest in time.
Longlora: Efficient fine-tuning of long-context large language models
Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia. 2023c · 2023
Closest in time.
Adapting Language Models to Compress Contexts
Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023 · 2023
Closest in time.
Dissecting transformer length extrapolation via the lens of receptive field analysis. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 13522–13537
Ta-Chung Chi, Ting-Han Fan, Alexander Rudnicky, and Peter Ramadge. 2023 · 2023
Closest in time.
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
Róbert Csordás, Piotr Piękos, and Kazuki Irie. 2023 · 2023
Closest in time.
How Long Can Open-Source LLMs Truly Promise on Context Length?
Anze Xie Ying Sheng Lianmin Zheng Joseph E. Gonzalez Ion Stoica Xuezhe Ma Dacheng Li*, Rulin Shao* and Hao Zhang. 2023 · 2023
Closest in time.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao. 2023 · 2023
Closest in time.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023 · 2023
Closest in time.
Longnet: Scaling transformers to 1,000,000,000 tokens
Jiayu Ding, Shuming Ma, Li Dong, Xingxing Zhang, Shaohan Huang, Wenhui Wang, and Furu Wei. 2023 · 2023
Closest in time.
A Survey on Long Text Modeling with Transformers
Zican Dong, Tianyi Tang, Lunyi Li, and Wayne Xin Zhao. 2023 · 2023
Closest in time.
Dynamically Scaled RoPE further increases performance of long context LLaMA with zero fine-tuning
emozilla. 2023 · 2023
Closest in time.
Extending Context Window of Large Language Models via Semantic Compression
Weizhi Fei, Xueyan Niu, Pingyi Zhou, Lu Hou, Bo Bai, Lei Deng, and Wei Han. 2023 · 2023
Closest in time.
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2023 · 2023
Closest in time.
OpenMoE: Open Mixture-of-Experts Language Models
Yao Fu Jinjie Ni Zangwei Zheng Wangchunshu Zhou Fuzhao Xue, Zian Zheng and Yang You. 2023 · 2023
Closest in time.
Llama/GPTNeoX: add RoPE scaling
gante. 2023 · 2023
Closest in time.
Lm-infinite: Simple on-the-fly length generalization for large language models
Chi Han, Qifan Wang, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang. 2023 · 2023
Closest in time.
Planning-oriented autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 17853–17862
Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, et al · 2023
Closest in time.
Efficient long-text understanding with short-text models
Maor Ivgi, Uri Shaham, and Jonathan Berant. 2023 · 2023
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Closest in time.
Llmlingua: Compressing prompts for accelerated inference of large language models
Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023b · 2023
Closest in time.
Introducing Qwen-7B: Open foundation and human-aligned models (of the state-of-the-arts)
JianxinMa. 2023 · 2023
Closest in time.
ChatGPT and large language model (LLM) chatbots: the current state of acceptability and a proposal for guidelines on utilization in academic medicine
Jin K Kim, Michael Chua, Mandy Rickard, and Armando Lorenzo. 2023 · 2023
Closest in time.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Closest in time.
RecallM: An Architecture for Temporal Context Understanding and Question Answering
Brandon Kynoch and Hugo Latapie. 2023 · 2023
Closest in time.
Prompted LLMs as Chatbot Modules for Long Open-domain Conversation
Gibbeum Lee, Volker Hartmann, Jongho Park, Dimitris Papailiopoulos, and Kangwook Lee. 2023 · 2023
Closest in time.
Accelerating distributed { \{ MoE } \} training and inference with lina. In 2023 USENIX Annual Technical Conference (USENIX ATC 23) . 945–959
Jiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang, and Hong Xu. 2023b · 2023
Closest in time.
Compressing context to enhance inference efficiency of large language models
Yucheng Li, Bo Dong, Chenghua Lin, and Frank Guerin. 2023a · 2023
Closest in time.
Loftq: Lora-fine-tuning-aware quantization for large language models
Yixiao Li, Yifan Yu, Chen Liang, Pengcheng He, Nikos Karampatziakis, Weizhu Chen, and Tuo Zhao. 2023d · 2023
Closest in time.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, Chuang Gan, and Song Han. 2023 · 2023
Closest in time.
Scaling Laws of RoPE-based Extrapolation
Xiaoran Liu, Hang Yan, Shuo Zhang, Chenxin An, Xipeng Qiu, and Dahua Lin. 2023 · 2023
Closest in time.
Large language models challenge the future of higher education
Silvia Milano, Joshua A McGrane, and Sabina Leonelli. 2023 · 2023
Closest in time.
RET-LLM: Towards a General Read-Write Memory for Large Language Models
Ali Modarressi, Ayyoob Imani, Mohsen Fayyaz, and Hinrich Schütze. 2023 · 2023
Closest in time.
Landmark Attention: Random-Access Infinite Context Length for Transformers
Amirkeivan Mohtashami and Martin Jaggi. 2023 · 2023
Closest in time.
FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement
Xiaonan Nie, Xupeng Miao, Zilong Wang, Zichao Yang, Jilong Xue, Lingxiao Ma, Gang Cao, and Bin Cui. 2023 · 2023
Closest in time.
OpenAI: GPT-4, 2023
OpenAI. 2023b · 2023
Closest in time.
Faster Causal Attention Over Large Sequences Through Sparse Flash Attention
Matteo Pagliardini, Daniele Paliotta, Martin Jaggi, and François Fleuret. 2023 · 2023
Closest in time.
Giraffe: Adventures in expanding context lengths in llms
Arka Pal, Deep Karkhanis, Manley Roberts, Samuel Dooley, Arvind Sundararajan, and Siddartha Naidu. 2023 · 2023
Closest in time.
Yarn: Efficient context window extension of large language models
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. 2023 · 2023
Closest in time.
Efficiently scaling transformer inference
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. 2023 · 2023
Closest in time.
From sparse to soft mixtures of experts
Joan Puigcerver, Carlos Riquelme, Basil Mustafa, and Neil Houlsby. 2023 · 2023
Closest in time.
PyTorch 2.0
Pytorch. 2023 · 2023
Closest in time.
Parallel Context Windows for Large Language Models
Nir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram, Inbal Magar, Omri Abend, Ehud Karpas, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023 · 2023
Closest in time.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Closest in time.
War of the chatbots: Bard, Bing Chat, ChatGPT, Ernie and beyond. The new AI gold rush and its impact on higher education
Jürgen Rudolph, Shannon Tan, and Samson Tan. 2023 · 2023
Closest in time.
Randomized Positional Encodings Boost Length Generalization of Transformers
Anian Ruoss, Grégoire Delétang, Tim Genewein, Jordi Grau-Moya, Róbert Csordás, Mehdi Bennani, Shane Legg, and Joel Veness. 2023 · 2023
Closest in time.
Memory Augmented Language Models through Mixture of Word Experts
Cicero Nogueira dos Santos, James Lee-Thorp, Isaac Noble, Chung-Ching Chang, and David Uthus. 2023 · 2023
Closest in time.
A Simple and Effective Pruning Approach for Large Language Models
Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter. 2023 · 2023
Closest in time.
A Frustratingly Easy Improvement for Position Embeddings via Random Padding
Mingxu Tao, Yansong Feng, and Dongyan Zhao. 2023 · 2023
Closest in time.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Closest in time.
Creating large language model applications utilizing langchain: A primer on developing llm apps fast. In International Conference on Applied Engineering and Natural Sciences , Vol. 1. 1050–1056
Oguzhan Topsakal and Tahir Cetin Akinci. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Focused Transformer: Contrastive Training for Context Scaling
Szymon Tworkowski, Konrad Staniszewski, Mikołaj Pacek, Yuhuai Wu, Henryk Michalewski, and Piotr Miłoś. 2023 · 2023
Closest in time.
Vanna: an open-source Python RAG framework for SQL generation and related functionality
Vanna-AI. 2023 · 2023
Closest in time.
ZeRO++: Extremely Efficient Collective Communication for Giant Model Training
Guanhua Wang, Heyang Qin, Sam Ade Jacobs, Connor Holmes, Samyam Rajbhandari, Olatunji Ruwase, Feng Yan, Lei Yang, and Yuxiong He. 2023b · 2023
Closest in time.
Augmenting Language Models with Long-Term Memory
Weizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, and Furu Wei. 2023a · 2023
Closest in time.
Using GitHub Copilot to solve simple programming problems. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 . 172–178
Michel Wermelinger. 2023 · 2023
Closest in time.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2023 · 2023
Closest in time.
Effective long-context scaling of foundation models
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, et al · 2023
Closest in time.
QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models
Yuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen, Heng Chang, Hengheng Zhang, Zhensu Chen, Xiaopeng Zhang, and Qi Tian. 2023 · 2023
Closest in time.
Burak Yetiştiren, Işık Özsoy, Miray Ayerdem, and Eray Tüzün. 2023 · 2023
Closest in time.
InfiniteBench: 128k Long-Context Benchmark for Language Models
Xinrong Zhang, Yingfa Chen, Shengding Hu, Qihao Wu, Junhao Chen, Zihang Xu, Zhenning Dai, Xu Han, Shuo Wang, Zhiyuan Liu, and Maosong Sun. 2023 · 2023
Closest in time.
Pytorch FSDP: experiences on scaling fully sharded data parallel
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al · 2023
Closest in time.
MemoryBank: Enhancing Large Language Models with Long-Term Memory
Wanjun Zhong, Lianghong Guo, Qiqi Gao, and Yanlin Wang. 2023 · 2023
Closest in time.
RecurrentGPT: Interactive Generation of (Arbitrarily) Long Text
Wangchunshu Zhou, Yuchen Eleanor Jiang, Peng Cui, Tiannan Wang, Zhenxin Xiao, Yifan Hou, Ryan Cotterell, and Mrinmaya Sachan. 2023 · 2023
Closest in time.
PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
Dawei Zhu, Nan Yang, Liang Wang, Yifan Song, Wenhao Wu, Furu Wei, and Sujian Li. 2023 · 2023
Closest in time.
LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
Hongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang, Zirui Liu, Chia-Yuan Chang, Huiyuan Chen, and Xia Hu. 2024 · 2024
Closest in time.
Memory-efficient fine-tuning of compressed large language models via sub-4-bit integer quantization
Jeonghoon Kim, Jung Hyun Lee, Sungdong Kim, Joonsuk Park, Kang Min Yoo, Se Jung Kwon, and Dongsoo Lee. 2024 · 2024
Closest in time.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024 · 2024
Closest in time.
Llm-pruner: On the structural pruning of large language models
Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2024 · 2024
Closest in time.
Learning to compress prompts with gist tokens
Jesse Mu, Xiang Li, and Noah Goodman. 2024 · 2024
Closest in time.
MemGPT: Towards LLMs as Operating Systems
Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. 2024 · 2024
Closest in time.
pipegoose: Large-scale 4D parallelism pre-training for ‘transformers‘
xrsrke. 2024 · 2024
Closest in time.