Understand
We propose "Areas of Attention", a novel attention-based model for automatic image captioning.
- Our approach models the dependencies between image regions, caption words, and the state of an RNN language model, using three pairwise interactions.
- In contrast to previous attention-based approaches that associate image regions only to the RNN state, our method allows a direct association between caption words and image regions.
- During training these associations are inferred from image-level captions, akin to weakly-supervised object detector training.
Built on
Nothing clear enough to list yet.
Similar
Nothing clear enough to list yet.
Then
Nothing clear enough to list yet.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…