Fetching the paper…

Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization · Around