Fetching the paper…

LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism · Around