Understand
Task-agnostic pre-training followed by task-specific fine-tuning is a default approach to train NLU models.
- Such models need to be deployed on devices across the cloud and the edge with varying resource and accuracy constraints.
- For a given task, repeating pre-training and fine-tuning across tens of devices is prohibitively expensive.
- We propose SuperShaper, a task agnostic pre-training approach which simultaneously pre-trains a large number of Transformer models by varying shapes, i.e., by varying the hidden dimensions across layers.
Built on
Nothing clear enough to list yet.
Similar
Nothing clear enough to list yet.
Then
Nothing clear enough to list yet.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…