Fetching the paper…

Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators · Around