Fetching the paper…
Reading the bibliography…
This study introduces bifurcated attention, a method designed to enhance language model inference in shared-context batch decoding scenarios.
Fast transformer decoding: One write-head is all you need
N. Shazeer · 1911
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2001
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Dissecting the NVIDIA volta GPU architecture via microbenchmarking
Z. Jia, M. Maggioni, B. Staiger, and D. P. Scarpazza · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
S. Dathathri, A. Madotto, J. Lan, J. Hung, E. Frank, P. Molino, J. Yosinski, and R. Liu · 2019
Earlier work this paper cites.
A study of bfloat16 for deep learning training
D. Kalamkar, D. Mudigere, N. Mellempudi, D. Das, K. Banerjee, S. Avancha, D. T. Vooturi, N. Jammalamadaka, J. Huang, H. Yuen, et al · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Z. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al · 2019
Earlier work this paper cites.
Simple and effective noisy channel modeling for neural machine translation
K. Yee, N. Ng, Y. N. Dauphin, and M. Auli · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
K. Guu, K. Lee, Z. Tung, P. Pasupat, and M. Chang · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2020
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models
M. Nadeem, A. Bethke, and S. Reddy · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
Unsupervised translation of programming languages
B. Roziere, M.-A. Lachaux, L. Chanussot, and G. Lample · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontañón, P. Pham, A. Ravula, Q. Wang, L. Yang, and A. Ahmed · 2020
Earlier work this paper cites.
Unified pre-training for program understanding and generation
W. U. Ahmad, S. Chakraborty, B. Ray, and K.-W. Chang · 2021
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. I. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. J. Cai, M. Terry, Q. V. Le, and C. Sutton · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba · 2021
Cited alongside, same era.
Nvidia a100 tensor core gpu: Performance and innovation
J. Choquette, W. Gandhi, O. Giroux, N. Stam, and R. Krashinsky · 2021
Cited alongside, same era.
Findings of the 2021 conference on machine translation (wmt21)
A. Farhad, A. Arkady, B. Magdalena, B. Ondřej, C. Rajen, C. Vishrav, M. R. Costa-jussa, E.-B. Cristina, F. Angela, F. Christian, et al · 2021
Cited alongside, same era.
Plug-and-blend: A framework for controllable story generation with blended control codes
Z. Lin and M. Riedl · 2021
Cited alongside, same era.
Wordcraft: story writing with large language models
A. Yuan, A. Coenen, E. Reif, and D. Ippolito · 2022
Later among the works it cites.
A survey on knowledge-enhanced pre-trained language models
C. Zhen, Y. Shang, X. Liu, Y. Li, Y. Chen, and D. Zhang · 2022
Later among the works it cites.
GQA: training generalized multi-query transformer models from multi-head checkpoints
J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebrón, and S. Sanghai · 2023
Later among the works it cites.
Santacoder: don’t reach for the stars!
L. B. Allal, R. Li, D. Kocetkov, C. Mou, C. Akiki, C. M. Ferrandis, N. Muennighoff, M. Mishra, A. Gu, M. Dey, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Cited alongside, same era.
Facebook ai wmt21 news translation task submission
C. Tran, S. Bhosale, J. Cross, P. Koehn, S. Edunov, and A. Fan · 2021
Cited alongside, same era.
Multi-lingual evaluation of code generation models
B. Athiwaratkun, S. K. Gouda, Z. Wang, X. Li, Y. Tian, M. Tan, W. U. Ahmad, S. Wang, Q. Sun, M. Shang, S. K. Gonugondla, H. Ding, V. Kumar, N. Fulton, A. Farahani, S. Jain, R. Giaquinto, H. Qian, M. K. Ramanathan, R. Nallapati, B. Ray, P. Bhatia, S. Sengupta, D. Roth, and B. Xiang · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
M. R. Costa-jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, et al · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Cited alongside, same era.
Llm.int8(): 8-bit matrix multiplication for transformers at scale
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh · 2022
Cited alongside, same era.
A. Bulatov, Y. Kuratov, and M. S. Burtsev · 2023
Later among the works it cites.
Accelerating large language model decoding with speculative sampling
C. Chen, S. Borgeaud, G. Irving, J. Lespiau, L. Sifre, and J. Jumper · 2023
Later among the works it cites.
Breaking the sequential dependency of llm inference using lookahead decoding, November 2023
Y. Fu, P. Bailis, I. Stoica, and H. Zhang · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention, 2023
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica · 2023
Later among the works it cites.
Starcoder: may the source be with you!
R. Li, L. B. Allal, Y. Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim, et al · 2023
Later among the works it cites.
Learning performance-improving code edits
A. Madaan, A. Shypula, U. Alon, M. Hashemi, P. Ranganathan, Y. Yang, G. Neubig, and A. Yazdanbakhsh · 2023
Later among the works it cites.
X. Miao, G. Oliaro, Z. Zhang, X. Cheng, Z. Wang, R. Y. Y. Wong, Z. Chen, D. Arfeen, R. Abhyankar, and Z. Jia · 2023
Later among the works it cites.
Co-writing screenplays and theatre scripts with language models: Evaluation by industry professionals
P. Mirowski, K. W. Mathewson, J. Pittman, and R. Evans · 2023
Later among the works it cites.
Codegen2: Lessons for training llms on programming and natural languages
E. Nijkamp, H. Hayashi, C. Xiong, S. Savarese, and Y. Zhou · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Accelerating generative AI with PyTorch II: GPT, Fast, 2023
T. PyTorch · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Later among the works it cites.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
M. N. Team · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models, 2023
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2023
Later among the works it cites.
Greener yet powerful: Taming large code generation models with quantization
X. Wei, S. K. Gonugondla, W. U. Ahmad, S. Wang, B. Ray, H. Qian, X. Li, V. Kumar, Z. Wang, Y. Tian, Q. Sun, B. Athiwaratkun, M. Shang, M. K. Ramanathan, P. Bhatia, and B. Xiang · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models, 2023
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan · 2023
Later among the works it cites.
Medusa: Simple llm inference acceleration framework with multiple decoding heads, 2024
T. Cai, Y. Li, Z. Geng, H. Peng, J. D. Lee, D. Chen, and T. Dao · 2024
Closest in time.
Hydragen: High-throughput llm inference with shared prefixes, 2024
J. Juravsky, B. Brown, R. Ehrlich, D. Y. Fu, C. Ré, and A. Mirhoseini · 2024
Closest in time.
Eagle: Speculative sampling requires rethinking feature uncertainty, 2024
Y. Li, F. Wei, C. Zhang, and H. Zhang · 2024
Closest in time.