Fetching the paper…
Reading the bibliography…
Nowadays, AI researchers become more and more interested in fine-tuning a pre-trained LLM, whose size has grown to up to over 100B parameters, for their downstream tasks.
T. Eicken, D. Culler, S. Goldstein, and K. Schauser, “Active messages: A mechanism for integrated communication and computation,” in ISCA , 1992
1992
Earlier work this paper cites.
I. Petrov, G. G. Almeida, A. P. Buchmann, and U. Gräf, “Building large storage based on flash disks,” in ADMS@VLDB , 2010
2010
Earlier work this paper cites.
NVIDIA, “Nvidia gpudirect: Enhancing data movement and access for gpus,” https://developer.nvidia.com/gpudirect
2011
Earlier work this paper cites.
Y. Lv, B. Cui, B. He, and X. Chen, “Operation-aware buffer management in flash-based systems,” in SIGMOD , 2011
2011
Earlier work this paper cites.
J. Do, D. Zhang, J. M. Patel, and D. J. DeWitt, “Fast peak-to-peak behavior with ssd buffer pool,” in ICDE , 2013
2013
Earlier work this paper cites.
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su, “Scaling distributed machine learning with the parameter server,” in OSDI , 2014
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint , 2014
2014
Earlier work this paper cites.
M. Rhu, N. Gimelshein, J. Clemons, A. Zulfiqar, and S. W. Keckler, “vdnn: Virtualized deep neural networks for scalable, memory-efficient neural network design,” in MICRO , 2016
2016
Earlier work this paper cites.
T. Chen, B. Xu, C. Zhang, and C. Guestrin, “Training deep nets with sublinear memory cost,” arXiv preprint , 2016
2016
Earlier work this paper cites.
J. Li, H.-W. Tseng, C. Lin, Y. Papakonstantinou, and S. Swanson, “Hippogriffdb: Balancing i/o and gpu bandwidth in big data analytics,” in VLDB , 2016
2016
Earlier work this paper cites.
M.-S. Kim, K. An, H. Park, H. Seo, and J. Kim, “Gts: A fast and scalable graph processing method based on streaming topology to gpus,” in SIGMOD , 2016
2016
Earlier work this paper cites.
A. Gruslys, R. Munos, I. Danihelka, M. Lanctot, and A. Graves, “Memory-efficient backpropagation through time,” in NIPS , 2016
2016
Earlier work this paper cites.
J. Jiang, B. Cui, C. Zhang, and L. Yu, “Heterogeneity-aware distributed parameter servers,” in SIGMOD , 2017
2017
Earlier work this paper cites.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in pytorch,” in NIPS Autodiff Workshop , 2017
2017
Earlier work this paper cites.
R. Thonangi and J. Yang, “On log-structured merge for solid-state drives,” in ICDE , 2017
2017
Earlier work this paper cites.
Y. Huang, T. Jin, Y. Wu, Z. Cai, X. Yan, F. Yang, J. Li, Y. Guo, and J. Cheng, “Flexps: Flexible parallelism control in parameter server architecture,” in VLDB , 2018
2018
Earlier work this paper cites.
J. Jiang, F. Fu, T. Yang, and B. Cui, “Sketchml: Accelerating distributed machine learning with data sketches,” in SIGMOD , 2018
2018
Earlier work this paper cites.
L. Wang, J. Ye, Y. Zhao, W. Wu, A. Li, S. L. Song, Z. Xu, and T. Kraska, “Superneurons: Dynamic gpu memory management for training deep neural networks,” in PPoPP , 2018
2018
Earlier work this paper cites.
T. D. Le, H. Imai, Y. Negishi, and K. Kawachiya, “Tflms: Large model support in tensorflow by graph rewriting,” arXiv preprint , 2018
2018
Earlier work this paper cites.
H. Jin, B. Liu, W. Jiang, Y. Ma, X. Shi, B. He, and S. Zhao, “Layer-centric memory reuse and data migration for extreme-scale deep learning on many-core architectures,” TACO , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT , 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI blog , 2019
2019
Earlier work this paper cites.
S. S. Sandha, W. Cabrera, M. Al-Kateb, S. Nair, and M. Srivastava, “In-database distributed machine learning: demonstration using teradata sql engine,” in VLDB , 2019
2019
Earlier work this paper cites.
Y. Chai, Y. Chai, X. Wang, H. Wei, N. Bao, and Y. Liang, “Ldc: a lower-level driven compaction method to optimize ssd-oriented key-value stores,” in ICDE , 2019
2019
Earlier work this paper cites.
M. Kusumoto, T. Inoue, G. Watanabe, T. Akiba, and M. Koyama, “A graph theoretic framework of recomputation algorithms for memory-efficient backpropagation,” in NeurIPS , 2019
2019
Earlier work this paper cites.
J. Zhang, S. H. Yeung, Y. Shu, B. He, and W. Wang, “Efficient memory management for gpu-based deep learning systems,” arXiv preprint , 2019
2019
Earlier work this paper cites.
S. Shriram, A. Garg, and P. Kulkarni, “Dynamic memory management for gpu-based training of deep neural networks,” in IPDPS , 2019
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in NeurIPS , 2020
2020
Earlier work this paper cites.
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro, “Megatron-lm: Training multi-billion parameter language models using model parallelism,” arXiv preprint , 2020
2020
Earlier work this paper cites.
N. Band, “Memflow: Memory-aware distributed deep learning,” in SIGMOD , 2020
2020
Earlier work this paper cites.
X. Peng, X. Shi, H. Dai, H. Jin, W. Ma, Q. Xiong, F. Yang, and X. Qian, “Capuchin: Tensor-based gpu memory management for deep learning,” in ASPLOS , 2020
2020
Earlier work this paper cites.
P. Jain, A. Jain, A. Nrusimha, A. Gholami, P. Abbeel, J. Gonzalez, K. Keutzer, and I. Stoica, “Checkmate: Breaking the memory wall with optimal tensor rematerialization,” in MLSys , 2020
2020
Earlier work this paper cites.
J. Chu, Y. Tu, Y. Zhang, and C. Weng, “Latte: A native table engine on nvme storage,” in ICDE , 2020
2020
Earlier work this paper cites.
M. Haubenschild, C. Sauer, T. Neumann, and V. Leis, “Rethinking logging, checkpoints, and recovery for high-performance storage engines,” in SIGMOD , 2020
2020
Earlier work this paper cites.
O. Beaumont, J. Herrmann, G. Pallez, and A. Shilova, “Optimal memory-aware backpropagation of deep join networks,” Philosophical Transactions of the Royal Society A , 2020
2020
Earlier work this paper cites.
M. Kirisame, S. Lyubomirsky, A. Haan, J. Brennan, M. He, J. Roesch, T. Chen, and Z. Tatlock, “Dynamic tensor rematerialization,” arXiv preprint , 2020
2020
Earlier work this paper cites.
O. Beaumont, L. Eyraud-Dubois, and A. Shilova, “Optimal gpu-cpu offloading strategies for deep neural network training,” in Euro-Par , 2020
2020
Earlier work this paper cites.
C.-C. Huang, G. Jin, and J. Li, “Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping,” in ASPLOS , 2020
2020
Earlier work this paper cites.
B. Pudipeddi, M. Mesmakhosroshahi, J. Xi, and S. Bharadwaj, “Training large neural networks with constant memory using a new execution algorithm,” arXiv preprint , 2020
2020
Earlier work this paper cites.
I. Trummer, “The case for nlp-enhanced database tuning: towards tuning tools that” read the manual”,” in VLDB , 2021
2021
Cited alongside, same era.
Y. Guo, Z. Zhang, J. Jiang, W. Wu, C. Zhang, B. Cui, and J. Li, “Model averaging in distributed machine learning: a case study with apache spark,” VLDBJ , 2021
2021
Cited alongside, same era.
X. Miao, X. Nie, Y. Shao, Z. Yang, J. Jiang, L. Ma, and B. Cui, “Heterogeneity-aware distributed machine learning training via partial reduce,” in SIGMOD , 2021
2021
Cited alongside, same era.
Y. Zhang, F. Mcquillan, N. Jayaram, N. Kak, E. Khanna, O. Kislal, D. Valdano, and A. Kumar, “Distributed deep learning on data systems: a comparative analysis of approaches,” in VLDB , 2021
2021
Cited alongside, same era.
K. Nagrecha, “Model-parallel model selection for deep learning systems,” in SIGMOD , 2021
2021
——, “Nvidia dgx platform,” https://www.nvidia.com/en-us/data-center/dgx-platform/
2023
Later among the works it cites.
T. Um, B. Oh, B. Seo, M. Kweun, G. Kim, and W.-Y. Lee, “Fastflow: Accelerating deep learning model training with smart offloading of input data pipeline,” in VLDB , 2023
2023
Later among the works it cites.
X. Miao, Y. Shi, Z. Yang, B. Cui, and Z. Jia, “Sdpipe: A semi-decentralized framework for heterogeneity-aware pipeline-parallel training,” in VLDB , 2023
2023
Later among the works it cites.
X. Nie, X. Miao, Z. Wang, Z. Yang, J. Xue, L. Ma, G. Cao, and B. Cui, “Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,” PACMMOD , 2023
2023
Later among the works it cites.
X. Nie, Y. Liu, F. Fu, J. Xue, D. Jiao, X. Miao, Y. Tao, and B. Cui, “Angel-ptm: A scalable and economical large-scale pre-training system in tencent,” in VLDB , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y. He, “Zero-offload: Democratizing billion-scale model training,” in ATC , 2021
2021
Cited alongside, same era.
J. Bae, J. Lee, Y. Jin, S. Son, S. Kim, H. Jang, T. J. Ham, and J. W. Lee, “Flashneuron: Ssd-enabled large-batch training of very deep neural networks,” in FAST , 2021
2021
Cited alongside, same era.
S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, and Y. He, “Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning,” in SC , 2021
2021
Cited alongside, same era.
A. Lerner and P. Bonnet, “Not your grandpa’s ssd: The era of co-designed storage devices,” in SIGMOD , 2021
2021
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV , 2021
2021
Cited alongside, same era.
M. D. M. Reddy, M. S. M. Basha, M. M. C. Hari, and M. N. Penchalaiah, “Dall-e: Creating images from text,” UGC Care Group I Journal , 2021
2021
Cited alongside, same era.
2023
Later among the works it cites.
Q. Zhou, H. Wang, X. Yu, C. Li, Y. Bai, F. Yan, and Y. Xu, “Mpress: Democratizing billion-scale model training on multi-gpu servers via memory-saving inter-operator parallelism,” in HPCA , 2023
2023
Later among the works it cites.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed, “Mistral 7b,” arXiv preprint , 2023
2023
Later among the works it cites.
Z. Bian, H. Liu, B. Wang, H. Huang, Y. Li, C. Wang, F. Cui, and Y. You, “Colossal-ai: A unified deep learning system for large-scale parallel training,” in ICPP , 2023
2023
Later among the works it cites.
G. Haas and V. Leis, “What modern nvme storage can do, and how to exploit it: High-performance i/o for high-performance storage engines,” in VLDB , 2023
2023
Later among the works it cites.
H. Zhang, Y. Zhou, Y. Xue, Y. Liu, and J. Huang, “G10: Enabling an efficient unified gpu memory and storage architecture with smart tensor migrations,” in MICRO , 2023
2023
Later among the works it cites.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in ICCV , 2023
2023
Later among the works it cites.
Y. Feng, M. Xie, Z. Tian, S. Wang, Y. Lu, and J. Shu, “Mobius: Fine tuning large-scale models on commodity gpu servers,” in ASPLOS , 2023
2023
Later among the works it cites.
Supermicro, “Supermicro sys-420gp-tnr dual xeon scalable 4u gpu superserver,” https://store.supermicro.com/us_en/4u-gpu-superserver-sys-420gp-tnr.html
2023
Later among the works it cites.
F. Maschi and G. Alonso, “The difficult balance between modern hardware and conventional cpus,” in DaMoN , 2023
2023
Later among the works it cites.
G. Alonso, N. Ailamaki, S. Krishnamurthy, S. Madden, S. Sivasubramanian, and R. Ramakrishnan, “Future of database system architectures,” in SIGMOD , 2023
2023
Later among the works it cites.
B. Lee, M. An, and S.-W. Lee, “Lru-c: Parallelizing database i/os for flash ssds,” in VLDB , 2023
2023
Later among the works it cites.
T. I. Papon and M. Athanassoulis, “Aceing the bufferpool management paradigm for modern storage devices,” in ICDE , 2023
2023
Later among the works it cites.
C. Duffy, J. Shim, S.-H. Kim, and J.-S. Kim, “Dotori: A key-value ssd based kv store,” in VLDB , 2023
2023
Later among the works it cites.
J. Sun, L. Su, Z. Shi, W. Shen, Z. Wang, L. Wang, J. Zhang, Y. Li, W. Yu, J. Zhou, and F. Wu, “Legion: Automatically pushing the envelope of Multi-GPU system for Billion-Scale GNN training,” in ATC , 2023
2023
Later among the works it cites.
Z. Zong, L. Lin, L. Lin, L. Wen, and Y. Sun, “Str: Hybrid tensor re-generation to break memory wall for dnn training,” TPDS , 2023
2023
Later among the works it cites.
L. Mei, K. Goetschalckx, A. Symons, and M. Verhelst, “Defines: Enabling fast exploration of the depth-first scheduling space for dnn accelerators through analytical modeling,” in HPCA , 2023
2023
Later among the works it cites.
J. Jung, J. Kim, and J. Lee, “Deepum: Tensor migration and prefetching in unified memory,” in ASPLOS , 2023
2023
Later among the works it cites.
LlamaTeam, “Thellama3herdofmodels,” https://ai.meta.com/research/publications/the-llama-3-herd-of-models/
2024
Closest in time.
X. Miao, Z. Jia, and B. Cui, “Demystifying data management for large language models,” in SIGMOD , 2024
2024
Closest in time.
C. Jin and S. Xie, “Fast-dit: Fast diffusion models with transformers,” https://github.com/chuanyangjin/fast-DiT
2024
Closest in time.
A. Kroviakov, P. Kurapov, C. Anneser, and J. Giceva, “Heterogeneous intra-pipeline device-parallel aggregations,” in DaMoN , 2024
2024
Closest in time.
A. Lerner and G. Alonso, “Data flow architectures for data processing on modern hardware,” in ICDE , 2024
2024
Closest in time.
Z. Wang, L. Shou, K. Chen, and X. Zhou, “Bushstore: Efficient b+ tree group indexing for lsm-tree in non-volatile memory,” in ICDE , 2024
2024
Closest in time.
X. Fan, S. Yan, Y. Huang, and C. Weng, “Tengine: A native distributed table storage engine,” in ICDE , 2024
2024
Closest in time.
Y. Wang, J. Yuan, S. Wu, H. Liu, J. Chen, C. Ma, and J. Qin, “Leaderkv: Improving read performance of kv stores via learned index and decoupled kv table,” in ICDE , 2024
2024
Closest in time.
Y. Wang, J. He, K. Sun, Y. Dong, J. Chen, C. Ma, A. C. Zhou, and R. Mao, “Boosting write performance of kv stores: An nvm-enabled storage collaboration approach,” in ICDE , 2024
2024
Closest in time.
Y. Huang, X. Fan, S. Yan, and C. Weng, “Neos: A nvme-gpus direct vector service buffer in user space,” in ICDE , 2024
2024
Closest in time.
T. I. Papon, “Enhancing data systems performance by exploiting ssd concurrency & asymmetry,” in ICDE , 2024
2024
Closest in time.
J. B. Park, V. S. Mailthody, Z. Qureshi, and W.-m. Hwu, “Accelerating sampling and aggregation operations in gnn frameworks with gpu initiated direct storage accesses,” in VLDB , 2024
2024
Closest in time.
M. Zhang, J. Sun, Q. Hu, P. Sun, Z. Wang, Y. Wen, and T. Zhang, “Torchgt: A holistic system for large-scale graph transformer training,” in SC , 2024
2024
Closest in time.
J. Herrmann, O. Beaumont, L. Eyraud-Dubois, J. Hermann, A. Joly, and A. Shilova, “Optimal checkpointing for heterogeneous chains: how to train deep neural networks with limited memory,” TOMS , 2024
2024
Closest in time.
J. Sun, Z. Shi, L. Su, W. Shen, Z. Wang, Y. Li, W. Yu, W. Lin, F. Wu, J. Zhou, and B. He, “Helios: Efficient distributed dynamic graph sampling for online gnn inference,” in PPoPP , 2025
2025
Closest in time.
J. Sun, M. Sun, Z. Zhang, J. Xie, Z. Shi, Z. Yang, J. Zhang, F. Wu, and Z. Wang, “Hyperion: Optimizing ssd access is all you need to enable cost-efficient out-of-core gnn training,” in ICDE , 2025
2025
Closest in time.