Fetching the paper…
Reading the bibliography…
Despite recent progress on computer vision and natural language processing, developing a machine that can understand video story is still hard to achieve due to the intrinsic difficulty of video story.
Constructing Hierarchical Q&A Datasets for Video Story Understanding
Heo, Y.; On, K.; Choi, S.; Lim, J.; Kim, J.; Ryu, J.; Bae, B.; and Zhang, B. 2019 · 1904
Earlier work this paper cites.
TVQA+: Spatio-Temporal Grounding for Video Question Answering
Lei, J.; Yu, L.; Berg, T. L.; and Bansal, M. 2019 · 1904
Earlier work this paper cites.
Psycho-logic: A possible alternative to Piaget’s formulation
Mclaughlin, G. H. 1963 · 1963
Earlier work this paper cites.
Cognitive development and cognitive style: A general psychological integration
Pascual-Leone, J. 1969 · 1969
Earlier work this paper cites.
Intellectual evolution from adolescence to adulthood
Piaget, J. 1972 · 1972
Earlier work this paper cites.
A Study of Concrete and Formal Operations in School Mathematics: A Piagetian Viewpoint
Collis, K. F. 1975 · 1975
Earlier work this paper cites.
Implications of neo-Piagetian theory for improving the design of instruction
Case, R. 1980 · 1980
Earlier work this paper cites.
Centering: A framework for modeling the local coherence of discourse
Grosz, B. J.; Weinstein, S.; and Joshi, A. K. 1995 · 1995
Earlier work this paper cites.
Interactive drama on computer: beyond linear narrative
Szilas, N. 1999 · 1999
Earlier work this paper cites.
Understanding script-based stories using commonsense reasoning
Mueller, E. T. 2004 · 2004
Earlier work this paper cites.
Narrative planning: Balancing plot and character
Riedl, M. O.; and Young, R. M. 2010 · 2010
Earlier work this paper cites.
A large video database for human motion recognition
Jhuang, H.; Garrote, H.; Poggio, E.; Serre, T.; and Hmdb, T. 2011 · 2011
Earlier work this paper cites.
Stanford Mobile Inquiry-based Learning Environment (SMILE): using mobile phones to promote student inquires in the elementary classroom
Seol, S.; Sharp, A.; and Kim, P. 2011 · 2011
Earlier work this paper cites.
The strong story hypothesis and the directed perception hypothesis
Winston, P. H. 2011 · 2011
Earlier work this paper cites.
A dataset of 101 human action classes from videos in the wild
Soomro, K.; Zamir, A. R.; and Shah, M. 2012 · 2012
Earlier work this paper cites.
Mctest: A challenge dataset for the open-domain machine comprehension of text
Richardson, M.; Burges, C. J.; and Renshaw, E. 2013 · 2013
Earlier work this paper cites.
Scripts, plans, goals, and understanding: An inquiry into human knowledge structures
Schank, R. C.; and Abelson, R. P. 2013 · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Karpathy, A.; Toderici, G.; Shetty, S.; Leung, T.; Sukthankar, R.; and Fei-Fei, L. 2014 · 2014
Cited alongside, same era.
GloVe: Global Vectors for Word Representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Cited alongside, same era.
Activitynet: A large-scale video benchmark for human activity understanding
Caba Heilbron, F.; Escorcia, V.; Ghanem, B.; and Carlos Niebles, J. 2015 · 2015
Cited alongside, same era.
Teaching machines to read and comprehend
Hermann, K. M.; Kocisky, T.; Grefenstette, E.; Espeholt, L.; Kay, W.; Suleyman, M.; and Blunsom, P. 2015 · 2015
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Cited alongside, same era.
End-To-End Memory Networks
The kinetics human action video dataset
Kay, W.; Carreira, J.; Simonyan, K.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; et al. 2017 · 2017
Later among the works it cites.
Dense-captioning events in videos
Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; and Carlos Niebles, J. 2017 · 2017
Later among the works it cites.
A Dataset and Exploration of Models for Understanding Video Data Through Fill-In-The-Blank Question-Answering
Maharaj, T.; Ballas, N.; Rohrbach, A.; Courville, A.; and Pal, C. 2017 · 2017
Later among the works it cites.
MarioQA: Answering Questions by Watching Gameplay Videos
Mun, J.; Seo, P. H.; Jung, I.; and Han, B. 2017 · 2017
Later among the works it cites.
Movie description
Rohrbach, A.; Torabi, A.; Rohrbach, M.; Tandon, N.; Pal, C.; Larochelle, H.; Courville, A.; and Schiele, B. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sukhbaatar, S.; szlam, a.; Weston, J.; and Fergus, R. 2015 · 2015
Cited alongside, same era.
Youtube-8m: A large-scale video classification benchmark
Abu-El-Haija, S.; Kothari, N.; Lee, J.; Natsev, P.; Toderici, G.; Varadarajan, B.; and Vijayanarasimhan, S. 2016 · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2015 · 2016
Cited alongside, same era.
The Goldilocks Principle: Reading Children’s Books with Explicit Memory Representations
Hill, F.; Bordes, A.; Chopra, S.; and Weston, J. 2016 · 2016
Cited alongside, same era.
A corpus and cloze evaluation for deeper understanding of commonsense stories
Mostafazadeh, N.; Chambers, N.; He, X.; Parikh, D.; Batra, D.; Vanderwende, L.; Kohli, P.; and Allen, J. 2016 · 2016
Cited alongside, same era.
A benchmark dataset and evaluation methodology for video object segmentation
Perazzi, F.; Pont-Tuset, J.; McWilliams, B.; Van Gool, L.; Gross, M.; and Sorkine-Hornung, A. 2016 · 2016
Cited alongside, same era.
Computational narrative intelligence: A human-centered goal for artificial intelligence
Riedl, M. O. 2016 · 2016
Cited alongside, same era.
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Later among the works it cites.
The narrativeqa reading comprehension challenge
Kočiskỳ, T.; Schwarz, J.; Blunsom, P.; Dyer, C.; Hermann, K. M.; Melis, G.; and Grefenstette, E. 2018 · 2018
Later among the works it cites.
TVQA: Localized, Compositional Video Question Answering
Lei, J.; Yu, L.; Bansal, M.; and Berg, T. L. 2018 · 2018
Later among the works it cites.
Structured Two-Stream Attention Network for Video Question Answering
Gao, L.; Zeng, P.; Song, J.; Li, Y.-F.; Liu, W.; Mei, T.; and Shen, H. T. 2019 · 2019
Later among the works it cites.
Beyond RNNs: Positional Self-Attention with Co-Attention for Video Question Answering
Li, X.; Song, J.; Gao, L.; Liu, X.; Huang, W.; Gan, C.; and He, X. 2019 · 2019
Later among the works it cites.
CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
Girdhar, R.; and Ramanan, D. 2020 · 2020
Closest in time.
Location-aware Graph Convolutional Networks for Video Question Answering
Huang, D.; Chen, P.; Zeng, R.; Du, Q.; Tan, M.; and Gan, C. 2020 · 2020
Closest in time.
Reasoning with Heterogeneous Graph Alignment for Video Question Answering
Jiang, P.; and Han, Y. 2020 · 2020
Closest in time.
Transformers: State-of-the-Art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Scao, T. L.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. M. 2020 · 2020
Closest in time.
CLEVRER: Collision Events for Video Representation and Reasoning
Yi, K.; Gan, C.; Li, Y.; Kohli, P.; Wu, J.; Torralba, A.; and Tenenbaum, J. B. 2020 · 2020
Closest in time.
Deepstory: Video story QA by deep embedded memory networks
Kim, K. M.; Heo, M. O.; Choi, S. H.; and Zhang, B. T. 2017 · 2022
Closest in time.