2021

SuperShaper: Task-Agnostic Super Pre-training of BERT Models with Variable Hidden Dimensions

Ganesan, Vinod, Ramesh, Gowtham, Kumar, Pratyush

Understand

Task-agnostic pre-training followed by task-specific fine-tuning is a default approach to train NLU models.

  • Such models need to be deployed on devices across the cloud and the edge with varying resource and accuracy constraints.
  • For a given task, repeating pre-training and fine-tuning across tens of devices is prohibitively expensive.
  • We propose SuperShaper, a task agnostic pre-training approach which simultaneously pre-trains a large number of Transformer models by varying shapes, i.e., by varying the hidden dimensions across layers.

Built on

Nothing clear enough to list yet.

Similar

Nothing clear enough to list yet.

Then

Nothing clear enough to list yet.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…