Fetching the paper…
Reading the bibliography…
AXLearn is a production system which facilitates scalable and high-performance training of large deep learning models.
Extensibility safety and performance in the spin operating system
Bershad, B. N., Savage, S., Pardyak, P., Sirer, E. G., Fiuczynski, M. E., Becker, D., Chambers, C., and Eggers, S · 1995
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning, 2014
Chetlur, S., Woolley, C., Vandermersch, P., Cohen, J., Tran, J., Catanzaro, B., and Shelhamer, E · 2014
Earlier work this paper cites.
Holistic configuration management at facebook
Tang, C., Kooburat, T., Venkatachalam, P., Chander, A., Wen, Z., Narayanan, A., Dowell, P., and Karl, R · 2015
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, M. X., Chen, D., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., and Chen, Z · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Zero: memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Efficient large-scale language model training on gpu clusters using megatron-lm
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., Phanishayee, A., and Zaharia, M · 2021
Earlier work this paper cites.
Gspmd: General and scalable parallelization for ml computation graphs, 2021
Xu, Y., Lee, H., Chen, D., Hechtman, B., Huang, Y., Joshi, R., Krikun, M., Lepikhin, D., Ly, A., Maggioni, M., Pang, R., Shazeer, N., Wang, S., Wang, T., Wu, Y., and Chen, Z · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Introducing chatgpt, 2022
OpenAI · 2022
Earlier work this paper cites.
URL https://github.com/google/orbax
Orbax Authors, 2022 · 2022
Cited alongside, same era.
Framework to configure and run machine learning experiments on top of jax., 2022
PAX team · 2022
Cited alongside, same era.
Orca: A distributed serving system for Transformer-Based generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Cited alongside, same era.
Alpa: Automating inter- and Intra-Operator parallelism for distributed deep learning
Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Xing, E. P., Gonzalez, J. E., and Stoica, I · 2022
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning, 2023
Dao, T · 2023
Cited alongside, same era.
Flax: A neural network library and ecosystem for JAX, 2023
DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Later among the works it cites.
URL https://openxla.org/xla/tf2xla
Xla: Optimizing compiler for machine learning., 2025 · 2025
Closest in time.
Neuron documentation, 2025
AWS · 2025
Closest in time.
The ai code editor, 2025
Cursor · 2025
Closest in time.
Gemini: Our most intelligent ai models, built for the agentic era, 2025
Google · 2025
Closest in time.
Haiku: Sonnet for jax, 2025
Haiku · 2025
Closest in time.
Maxtext., 2025
MaxText team · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Heek, J., Levskaya, A., Oliver, A., Ritter, M., Rondepierre, B., Steiner, A., and van Zee, M · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Sequence parallelism: Long sequence training from system perspective
Li, S., Xue, F., Baranwal, C., Li, Y., and You, Y · 2023
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2023
Cited alongside, same era.
Pytorch fsdp: Experiences on scaling fully sharded data parallel
Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Damania, P., Nguyen, B., Chauhan, G., Hao, Y., Mathews, A., and Li, S · 2023
Cited alongside, same era.
Revisiting moe and dense speed-accuracy comparisons for llm training, 2024
Du, X., Gunter, T., Kong, X., Lee, M., Wang, Z., Zhang, A., Du, N., and Pang, R · 2024
Cited alongside, same era.
MegaScale: Scaling large language model training to more than 10,000 GPUs
Jiang, Z., Lin, H., Zhong, Y., Huang, Q., Chen, Y., Zhang, Z., Peng, Y., Li, X., Xie, C., Nong, S., Jia, Y., He, S., Chen, H., Bai, Z., Hou, Q., Yan, S., Zhou, D., Sheng, Y., Jiang, Z., Xu, H., Wei, H., Zhang, Z., Nie, P., Zou, L., Zhao, S., Xiang, L., Liu, Z., Li, Z., Jia, X., Ye, J., Jin, X., and Liu, X · 2024
Cited alongside, same era.
Closest in time.
URL https://github.com/google-deepmind/optax
Optax Authors, 2025 · 2025
Closest in time.
Pallas: a jax kernel language., 2025
pallas · 2025
Closest in time.
Fully sharded data parallel in pytorch xla, 2025
PyTorch · 2025
Closest in time.
Operation semantics - openxla project, 2025
XLA · 2025
Closest in time.
Work happy with zoom ai companion, 2025
Zoom · 2025
Closest in time.
The click modular router
Kohler, E., Morris, R., Chen, B., Jannotti, J., and Kaashoek, M. F · 2071
Closest in time.