Fetching the paper…
Reading the bibliography…
We first propose a new task named Dialogue Description (Dial2Desc).
Nothing clear enough to list yet.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K.; Ba, J.; Kiros, R.; Cho, K.; Courville, A. C.; Salakhutdinov, R.; Zemel, R. S.; and Bengio, Y
Cited in the paper.
Nothing clear enough to list yet.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…