Fetching the paper…

Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos · Around