Fetching the paper…

Grass: Compute Efficient Low-Memory LLM Training with Structured Sparse Gradients · Around