Fetching the paper…

Dynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled Transformers · Around