Understand
Visual question answering (VQA) task not only bridges the gap between images and language, but also requires that specific contents within the image are understood as indicated by linguistic context of the question, in order to generate the accurate answers.
- Thus, it is critical to build an efficient embedding of images and texts.
- We implement DualNet, which fully takes advantage of discriminative power of both image and textual features by separately performing two operations.
- Building an ensemble of DualNet further boosts the performance.
Built on
Nothing clear enough to list yet.
Similar
Nothing clear enough to list yet.
Then
Nothing clear enough to list yet.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…