Fetching the paper…
Reading the bibliography…
Videos of actions are complex signals containing rich compositional structure in space and time.
The lexical nature of syntactic ambiguity resolution
MacDonald, M. C., Pearlmutter, N. J., and Seidenberg, M. S · 1994
Earlier work this paper cites.
Learning subjective language
Wiebe, J., Wilson, T., Bruce, R., Bell, M., and Martin, M · 2004
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Image retrieval using scene graphs
Johnson, J., Krishna, R., Stark, M., Li, L.-J., Shamma, D., Bernstein, M., and Fei-Fei, L · 2015
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Mathieu, M., Couprie, C., and LeCun, Y · 2015
Earlier work this paper cites.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Schuster, S., Krishna, R., Chang, A., Fei-Fei, L., and Manning, C. D · 2015
Earlier work this paper cites.
Interaction networks for learning about objects, relations and physics
Battaglia, P., Pascanu, R., Lai, M., Rezende, D. J., et al · 2016
Earlier work this paper cites.
Structural-rnn: Deep learning on spatio-temporal graphs
Jain, A., Zamir, A. R., Savarese, S., and Saxena, A · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Kipf, T. N. and Welling, M · 2016
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric
Larsen, A. B. L., Sønderby, S. K., Larochelle, H., and Winther, O · 2016
Earlier work this paper cites.
Improved techniques for training gans
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., Chen, X., and Chen, X · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Sigurdsson, G. A., Varol, G., Wang, X., Farhadi, A., Laptev, I., and Gupta, A · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
Vondrick, C., Pirsiavash, H., and Torralba, A · 2016
Earlier work this paper cites.
An uncertain future: Forecasting from static images using variational autoencoders
Walker, J., Doersch, C., Gupta, A., and Hebert, M · 2016
Earlier work this paper cites.
View synthesis by appearance flow
Zhou, T., Tulsiani, S., Sun, W., Malik, J., and Efros, A. A · 2016
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
Goyal, R., Kahou, S. E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Flownet 2.0: Evolution of optical flow estimation with deep networks
Ilg, E., Mayer, N., Saikia, T., Keuper, M., Dosovitskiy, A., and Brox, T · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Isola, P., Zhu, J.-Y., Zhou, T., and Efros, A. A · 2017
Earlier work this paper cites.
Attentive semantic video generation using captions
Marwah, T., Mittal, G., and Balasubramanian, V. N · 2017
Earlier work this paper cites.
Pixels to graphs by associative embedding
Newell, A. and Deng, J · 2017
Earlier work this paper cites.
To create what you tell: Generating videos from captions
Pan, Y., Qiu, Z., Yao, T., Li, H., and Mei, T · 2017
Earlier work this paper cites.
Temporal generative adversarial nets with singular value clipping
Saito, M., Matsumoto, E., and Saito, S · 2017
Cited alongside, same era.
Visual interaction networks: Learning a physics simulator from video
Watters, N., Zoran, D., Weber, T., Battaglia, P., Pascanu, R., and Tacchetti, A · 2017
Cited alongside, same era.
Scene Graph Generation by Iterative Message Passing
Xu, D., Zhu, Y., Choy, C. B., and Fei-Fei, L · 2017
Cited alongside, same era.
Stochastic variational video prediction
Babaeizadeh, M., Finn, C., Erhan, D., Campbell, R. H., and Levine, S · 2018
Cited alongside, same era.
Recycle-gan: Unsupervised video retargeting
Bansal, A., Ma, S., Ramanan, D., and Sheikh, Y · 2018
Cited alongside, same era.
Probabilistic neural programmed networks for scene generation
Deng, Z., Chen, J., Fu, Y., and Mori, G · 2018
Cited alongside, same era.
Everybody dance now
Chan, C., Ginosar, S., Zhou, T., and Efros, A. A · 2019
Later among the works it cites.
Text-based editing of talking-head video
Fried, O., Tewari, A., Zollhöfer, M., Finkelstein, A., Shechtman, E., Goldman, D. B., Genova, K., Jin, Z., Theobalt, C., and Agrawala, M · 2019
Later among the works it cites.
Learning individual styles of conversational gesture
Ginosar, S., Bar, A., Kohavi, G., Chan, C., Owens, A., and Malik, J · 2019
Later among the works it cites.
Video action transformer network
Girdhar, R., Carreira, J., Doersch, C., and Zisserman, A · 2019
Later among the works it cites.
Spatio-temporal action graph networks
Herzig, R., Levi, E., Xu, H., Gao, H., Brosh, E., Wang, X., Globerson, A., and Darrell, T · 2019
Later among the works it cites.
Action genome: Actions as composition of spatio-temporal scene graphs
Ji, J., Krishna, R., Fei-Fei, L., and Niebles, J. C · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic video generation with a learned prior
Denton, E. and Fergus, R · 2018
Cited alongside, same era.
Imagine this! scripts to compositions to videos
Gupta, T., Schwenk, D., Farhadi, A., Hoiem, D., and Kembhavi, A · 2018
Cited alongside, same era.
Mapping images to scene graphs with permutation-invariant structured prediction
Herzig, R., Raboh, M., Chechik, G., Berant, J., and Globerson, A · 2018
Cited alongside, same era.
Image generation from scene graphs
Johnson, J., Gupta, A., and Fei-Fei, L · 2018
Cited alongside, same era.
Neural relational inference for interacting systems
Kipf, T., Fetaya, E., Wang, K.-C., Welling, M., and Zemel, R · 2018
Cited alongside, same era.
Referring relationships
Krishna, R., Chami, I., Bernstein, M. S., and Fei-Fei, L · 2018
Cited alongside, same era.
Later among the works it cites.
Deep video inpainting
Kim, D., Woo, S., Lee, J.-Y., and Kweon, I. S · 2019
Later among the works it cites.
Tsm: Temporal shift module for efficient video understanding
Lin, J., Gan, C., and Han, S · 2019
Later among the works it cites.
Semantic image synthesis with spatially-adaptive normalization
Park, T., Liu, M.-Y., Wang, T.-C., and Zhu, J.-Y · 2019
Later among the works it cites.
Triplet-aware scene graph embeddings
Schroeder, B., Tripathi, S., and Tang, H · 2019
Later among the works it cites.
Animating arbitrary objects via deep motion transfer
Siarohin, A., Lathuilière, S., Tulyakov, S., Ricci, E., and Sebe, N · 2019
Later among the works it cites.
High fidelity video prediction with large stochastic recurrent neural networks
Villegas, R., Pathak, A., Kannan, H., Erhan, D., Le, Q. V., and Lee, H · 2019
Later among the works it cites.
Few-shot video-to-video synthesis
Wang, T.-C., Liu, M.-Y., Tao, A., Liu, G., Kautz, J., and Catanzaro, B · 2019
Later among the works it cites.
Scene graph captioner: Image captioning based on structural visual representation
Xu, N., Liu, A.-A., Liu, J., Nie, W., and Su, Y · 2019
Later among the works it cites.
Compositional video prediction
Ye, Y., Singh, M., Gupta, A., and Tulsiani, S · 2019
Later among the works it cites.
Clevrer: Collision events for video representation and reasoning
Yi, K., Gan, C., Li, Y., Kohli, P., Wu, J., Torralba, A., and Tenenbaum, J. B · 2019
Later among the works it cites.
CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
Girdhar, R. and Ramanan, D · 2020
Closest in time.
Learning canonical representations for scene graph to image generation
Herzig, R., Bar, A., Xu, H., Chechik, G., Darrell, T., and Globerson, A · 2020
Closest in time.
Analyzing and improving the image quality of stylegan
Karras, T., Laine, S., Aittala, M., Hellsten, J., Lehtinen, J., and Aila, T · 2020
Closest in time.
Videoflow: A conditional flow-based model for stochastic video generation
Kumar, M., Babaeizadeh, M., Erhan, D., Finn, C., Levine, S., Dinh, L., and Kingma, D · 2020
Closest in time.
World-consistent video-to-video synthesis
Mallya, A., Wang, T.-C., Sapra, K., and Liu, M.-Y · 2020
Closest in time.
Something-else: Compositional action recognition with spatial-temporal interaction networks
Materzynska, J., Xiao, T., Herzig, R., Xu, H., Wang, X., and Darrell, T · 2020
Closest in time.
Differentiable scene graphs
Raboh, M., Herzig, R., Chechik, G., Berant, J., and Globerson, A · 2020
Closest in time.