Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are a new and powerful tool for a wide span of applications involving natural language and demonstrate impressive code generation abilities.
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, A. M. Rush, Huggingface’s transformers: State-of-the-art natural language processing (2020) · 1910
Earlier work this paper cites.
S. Rajbhandari, J. Rasley, O. Ruwase, Y. He, Zero: Memory optimizations toward training trillion parameter models (2020) · 1910
Earlier work this paper cites.
doi:10.3115/992424.992434
J. J. Webster, C. Kit, Tokenization as the initial phase in nlp , in: Proceedings of the 14th Conference on Computational Linguistics - Volume 4, COLING ’92, Association for Computational Linguistics, USA, 1992, p. 1106–1110 · 1992
Earlier work this paper cites.
doi:10.1145/780822.781156
S. Lerner, T. Millstein, C. Chambers, Automatically proving the correctness of compiler optimizations , SIGPLAN Not. 38 (5) (2003) 220–231 · 2003
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, D. Amodei, Language models are few-shot learners (2020) · 2005
Earlier work this paper cites.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. tau Yih, T. Rocktäschel, S. Riedel, D. Kiela, Retrieval-augmented generation for knowledge-intensive nlp tasks (2021) · 2005
Earlier work this paper cites.
doi:10.1145/1993498.1993532
X. Yang, Y. Chen, E. Eide, J. Regehr, Finding and understanding bugs in c compilers , in: Proceedings of the 32nd ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI ’11, Association for Computing Machinery, New York, NY, USA, 2011, p. 283–294 · 2011
Earlier work this paper cites.
W. Sawyer, G. Zaengl, L. Linardakis, Towards a multi-node openacc implementation of the icon model, in: EGU General Assembly Conference Abstracts, 2014, p. 15276
2014
Earlier work this paper cites.
X. Lapillonne, O. Fuhrer, Using compiler directives to port large scientific applications to gpus: An example from atmospheric science, Parallel Processing Letters 24 (01) (2014) 1450003
2014
Earlier work this paper cites.
doi:10.1145/2594291.2594334
V. Le, M. Afshari, Z. Su, Compiler validation via equivalence modulo inputs , in: Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI ’14, Association for Computing Machinery, New York, NY, USA, 2014, p. 216–226 · 2014
Earlier work this paper cites.
S. Sathe, Accelerating the ansys fluent r18. 0 radiation solver with openacc (2016)
2016
Earlier work this paper cites.
R. Sennrich, B. Haddow, A. Birch, Neural machine translation of rare words with subword units (2016) · 2016
Earlier work this paper cites.
NVIDIA, CUDA SDK Code Samples, http://developer.nvidia.com/cuda-cc-sdk-code-samples , accessed: 2017-02-03
2017
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, D. Amodei, Deep reinforcement learning from human preference?s , in: I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, R. Garnett (Eds.), Advances in Neural Information Processing Systems, Vol. 30, Curran Associates, Inc., 2017, p. 1. URL https://proceedings.neurips.cc/paper_files/paper/2017/file/d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf
2017
Earlier work this paper cites.
J. E. Denny, S. Lee, J. S. Vetter, Clacc: Translating openacc to openmp in clang, in: 2018 IEEE/ACM 5th Workshop on the LLVM Compiler Infrastructure in HPC (LLVM-HPC), IEEE, 2018, pp. 18–29
2018
Earlier work this paper cites.
R. Searles, S. Chandrasekaran, W. Joubert, O. Hernandez, Mpi+ openacc: Accelerating radiation transport mini-application, minisweep, on heterogeneous systems, Computer Physics Communications 236 (2019) 176–187
2019
Earlier work this paper cites.
A. Alpay, V. Heuveline, Sycl beyond opencl: The architecture, current state and future direction of hipsycl, in: Proceedings of the International Workshop on OpenCL, 2020, pp. 1–1
2020
Earlier work this paper cites.
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, S. Huang, Trl: Transformer reinforcement learning, https://github.com/huggingface/trl (2020)
2020
Earlier work this paper cites.
Anaconda software distribution (2020). URL https://docs.anaconda.com/
2020
Earlier work this paper cites.
B. Kim, K. S. Yoon, H.-J. Kim, Gpu-accelerated laplace equation model development based on cuda fortran, Water 13 (23) (2021) 3435
2021
Earlier work this paper cites.
A. Tamkin, M. Brundage, J. Clark, D. Ganguli, Understanding the capabilities, limitations, and societal impact of large language models (2021) · 2021
Earlier work this paper cites.
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, W. Zaremba, Evaluating large language models trained on code (2021) · 2021
Earlier work this paper cites.
B. Lester, R. Al-Rfou, N. Constant, The power of scale for parameter-efficient prompt tuning (2021) · 2021
Earlier work this paper cites.
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, C. Sutton, Program synthesis with large language models (2021) · 2021
Earlier work this paper cites.
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, J. Schulman, Training verifiers to solve math word problems (2021) · 2021
Cited alongside, same era.
M. Tufano, D. Drain, A. Svyatkovskiy, S. K. Deng, N. Sundaresan, Unit test case generation with transformers and focal context , arXiv (May 2021). URL https://www.microsoft.com/en-us/research/publication/unit-test-case-generation-with-transformers-and-focal-context/
2021
Cited alongside, same era.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, Lora: Low-rank adaptation of large language models (2021) · 2021
Cited alongside, same era.
codeium, Codeium · free ai code completion & chat (2022). URL https://codeium.com/
2022
Cited alongside, same era.
V. Zouhar, C. Meister, J. L. Gastaldi, L. Du, T. Vieira, M. Sachan, R. Cotterell, A formal perspective on byte-pair encoding (2023) · 2023
Closest in time.
J. Liu, C. S. Xia, Y. Wang, L. ZHANG, Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation , in: Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=1qvx610Cu7
2023
Closest in time.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, D. Zhou, Chain-of-thought prompting elicits reasoning in large language models (2023) · 2023
Closest in time.
T. Dettmers, A. Pagnoni, A. Holtzman, L. Zettlemoyer, Qlora: Efficient finetuning of quantized llms (2023) · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
M. Stack, P. Macklin, R. Searles, S. Chandrasekaran, Openacc acceleration of an agent-based biological simulation framework, Computing in Science & Engineering 24 (5) (2022) 53–63
2022
Cited alongside, same era.
doi:10.1109/WACCPD56842.2022.00006
A. Jarmusch, A. Liu, C. Munley, D. Horta, V. Ravichandran, J. Denny, K. Friedline, S. Chandrasekaran, Analysis of validating and verifying openacc compilers 3.0 and above, in: 2022 Workshop on Accelerator Programming Using Directives (WACCPD), 2022, pp. 1–10 · 2022
Cited alongside, same era.
T. Huber, S. Pophale, N. Baker, M. Carr, N. Rao, J. Reap, K. Holsapple, J. H. Davis, T. Burnus, S. Lee, et al., Ecp sollve: Validation and verification testsuite status update and compiler insight for openmp, in: 2022 IEEE/ACM International Workshop on Performance, Portability and Productivity in HPC (P3HPC), IEEE, 2022, pp. 123–135
2022
Cited alongside, same era.
Anthropic, Introucing claude (2022). URL https://www.anthropic.com/index/introducing-claude
2022
Cited alongside, same era.
J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, Q. V. Le, Finetuned language models are zero-shot learners , in: International Conference on Learning Representations, 2022, p. 1. URL https://openreview.net/forum?id=gEZrGCozdqR
2022
Cited alongside, same era.
N. Lambert, L. Castricato, L. V. Werra, A. Havrilla, Illustrating reinforcement learning from human feedback (rlhf), https://huggingface.co/blog/rlhf (2022)
2022
Cited alongside, same era.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, R. Lowe, Training language models to follow instructions with human feedback (2022) · 2022
Cited alongside, same era.
AWS, Introducing amazon codewhisperer, the ml-powered coding companion (2023). URL https://aws.amazon.com/blogs/machine-learning/introducing-amazon-codewhisperer-the-ml-powered-coding-companion/
2023
Closest in time.
R. Li, L. B. Allal, Y. Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim, Q. Liu, E. Zheltonozhskii, T. Y. Zhuo, T. Wang, O. Dehaene, M. Davaadorj, J. Lamy-Poirier, J. Monteiro, O. Shliazhko, N. Gontier, N. Meade, A. Zebaze, M.-H. Yee, L. K. Umapathi, J. Zhu, B. Lipkin, M. Oblokulov, Z. Wang, R. Murthy, J. Stillerman, S. S. Patel, D. Abulkhanov, M. Zocca, M. Dey, Z. Zhang, N. Fahmy, U. Bhattacharyya, W. Yu, S. Singh, S. Luccioni, P. Villegas, M. Kunakov, F. Zhdanov, M. Romero, T. Lee, N. Timor, J. Ding, C. Schlesinger, H. Schoelkopf, J. Ebert, T. Dao, M. Mishra, A. Gu, J. Robinson, C. J. Anderson, B. Dolan-Gavitt, D. Contractor, S. Reddy, D. Fried, D. Bahdanau, Y. Jernite, C. M. Ferrandis, S. Hughes, T. Wolf, A. Guha, L. von Werra, H. de Vries, Starcoder: may the source be with you! (2023) · 2023
Closest in time.
Z. Luo, C. Xu, P. Zhao, Q. Sun, X. Geng, W. Hu, C. Tao, J. Ma, Q. Lin, D. Jiang, Wizardcoder: Empowering code large language models with evol-instruct (2023) · 2023
Closest in time.
L. Chen, P.-H. Lin, T. Vanderbruggen, C. Liao, M. Emani, B. de Supinski, Lm4hpc: Towards effective language model application in high-performance computing (2023) · 2023
Closest in time.
T. Kadosh, N. Hasabnis, V. A. Vo, N. Schneider, N. Krien, A. Wasay, N. Ahmed, T. Willke, G. Tamir, Y. Pinter, T. Mattson, G. Oren, Scope is all you need: Transforming llms for hpc code (2023) · 2023
Closest in time.
doi:10.1145/3605731.3605886
W. Godoy, P. Valero-Lara, K. Teranishi, P. Balaprakash, J. Vetter, Evaluation of openai codex for hpc parallel programming models kernel generation , in: Proceedings of the 52nd International Conference on Parallel Processing Workshops, ICPP Workshops ’23, Association for Computing Machinery, New York, NY, USA, 2023, p. 136–144 · 2023
Closest in time.
M. Schäfer, S. Nadi, A. Eghbali, F. Tip, An empirical evaluation of using large language models for automated unit test generation (2023) · 2023
Closest in time.
Phind, Phind/phind-codellama-34b-v2 · hugging face (2023). URL https://huggingface.co/Phind/Phind-CodeLlama-34B-v2
2023
Closest in time.
N. McKenna, T. Li, L. Cheng, M. J. Hosseini, M. Johnson, M. Steedman, Sources of hallucination by large language models on inference tasks (2023) · 2023
Closest in time.
facebookresearch, codellama (2023). URL https://github.com/facebookresearch/codellama
2023
Closest in time.
Langchain, Langchain (2023). URL https://www.langchain.com/
2023
Closest in time.
OpenAI, Openai api (2023). URL https://openai.com/blog/openai-api
2023
Closest in time.
T. Dao, Flashattention-2: Faster attention with better parallelism and work partitioning (2023) · 2023
Closest in time.
Y. Zhao, A. Gu, R. Varma, L. Luo, C.-C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y. Hao, A. Mathews, S. Li, Pytorch fsdp: Experiences on scaling fully sharded data parallel (2023) · 2023
Closest in time.
NVIDIA, Nvidia hpc sdk (2023). URL https://developer.nvidia.com/hpc-sdk
2023
Closest in time.
NERSC, Using perlmutter (2023). URL https://docs.nersc.gov/systems/perlmutter/
2023
Closest in time.
U. of Delaware, Darwin (2023). URL https://dsi.udel.edu/core/computational-resources/darwin/
2023
Closest in time.
D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. K. Li, F. Luo, Y. Xiong, W. Liang, Deepseek-coder: When the large language model meets programming – the rise of code intelligence (2024) · 2024
Closest in time.
doi:10.3390/app12178805
M. Mars, From word embeddings to pre-trained language models: A state-of-the-art walkthrough , Applied Sciences 12 (17) · 2076
Closest in time.