Fetching the paper…

Self-supervised video pretraining yields robust and more human-aligned visual representations · Around