Fetching the paper…
Reading the bibliography…
Transformers are slow to train on videos due to extremely large numbers of input tokens, even though many video tokens are repeated over time.
Nothing clear enough to list yet.
Nothing clear enough to list yet.
Nothing clear enough to list yet.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…