Fetching the paper…

Cross modal video representations for weakly supervised active speaker localization · Around