Fetching the paper…
Reading the bibliography…
Automatic generation of textual video descriptions that are time-aligned with video content is a long-standing goal in computer vision.
Nothing clear enough to list yet.
Nothing clear enough to list yet.
Nothing clear enough to list yet.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…