Fetching the paper…
Reading the bibliography…
In this study, we aim to predict the plausible future action steps given an observation of the past and study the task of instructional activity anticipation.
Y. Li, J. Yang, Y. Song, L. Cao, J. Luo, and L.-J. Li, “Learning from noisy labels with distillation,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 1910–1918
1918
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
M. S. Ryoo, “Human activity prediction: Early recognition of ongoing activities from streaming videos,” in 2011 International Conference on Computer Vision . IEEE, 2011, pp. 1036–1043
2011
Earlier work this paper cites.
M. Hoai and F. De la Torre, “Max-margin early event detectors,” International Journal of Computer Vision , vol. 107, no. 2, pp. 191–202, 2014
2014
Earlier work this paper cites.
T. Lan, T.-C. Chen, and S. Savarese, “A hierarchical representation for future action prediction,” in European Conference on Computer Vision . Springer, 2014, pp. 689–704
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Z. Xu, L. Qing, and J. Miao, “Activity auto-completion: Predicting human activities from partial videos,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 3191–3199
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko, “Sequence to sequence-video to text,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 4534–4542
2015
Earlier work this paper cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3128–3137
2015
Earlier work this paper cites.
C. Vondrick, H. Pirsiavash, and A. Torralba, “Anticipating visual representations from unlabeled video,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 98–106
2016
Earlier work this paper cites.
A. Alahi, K. Goel, V. Ramanathan, A. Robicquet, L. Fei-Fei, and S. Savarese, “Social lstm: Human trajectory prediction in crowded spaces,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 961–971
2016
Earlier work this paper cites.
S. Gupta, J. Hoffman, and J. Malik, “Cross modal distillation for supervision transfer,” in CVPR , 2016, pp. 2827–2836
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
W.-C. Ma, D.-A. Huang, N. Lee, and K. M. Kitani, “Forecasting interactive dynamics of pedestrians with fictitious play,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 774–782
2017
Cited alongside, same era.
S. Zagoruyko and N. Komodakis, “Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer,” in ICLR , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
K. Li, Y. Zhang, K. Li, Y. Li, and Y. Fu, “Visual semantic reasoning for image-text matching,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 4654–4662
2019
Later among the works it cites.
C.-Y. Chang, D.-A. Huang, D. Xu, E. Adeli, L. Fei-Fei, and J. C. Niebles, “Procedure planning in instructional videos,” in European Conference on Computer Vision . Springer, 2020, pp. 334–350
2020
Later among the works it cites.
J. Ji, R. Krishna, L. Fei-Fei, and J. C. Niebles, “Action genome: Actions as compositions of spatio-temporal scene graphs,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 10 236–10 247
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Salvador, N. Hynes, Y. Aytar, J. Marin, F. Ofli, I. Weber, and A. Torralba, “Learning cross-modal embeddings for cooking recipes and food images,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 3020–3028
2017
Cited alongside, same era.
Y. Abu Farha, A. Richard, and J. Gall, “When will you do what?-anticipating temporal occurrences of activities,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 5343–5352
2018
Cited alongside, same era.
C. Rodriguez, B. Fernando, and H. Li, “Action anticipation by predicting future dynamic images,” in Proceedings of the European Conference on Computer Vision (ECCV) Workshops , 2018, pp. 0–0
2018
Cited alongside, same era.
J. Mun, K. Lee, J. Shin, and B. Han, “Learning to specialize with knowledge distillation for visual question answering,” in Advances in Neural Information Processing Systems , 2018, pp. 8081–8091
2018
Cited alongside, same era.
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, “End-to-end dense video captioning with masked transformer,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 8739–8748
2018
Cited alongside, same era.
F. Sener and A. Yao, “Zero-shot anticipation for instructional activities,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 862–871
2019
Cited alongside, same era.
Q. Ke, M. Fritz, and B. Schiele, “Time-conditioned action anticipation in one shot,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 9925–9934
2019
Cited alongside, same era.
M.-F. Chang, J. Lambert, P. Sangkloy, J. Singh, S. Bak, A. Hartnett, D. Wang, P. Carr, S. Lucey, D. Ramanan et al. , “Argoverse: 3d tracking and forecasting with rich maps,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 8748–8757
2019
Cited alongside, same era.
Y. B. Ng and B. Fernando, “Forecasting future action sequences with attention: a new approach to weakly supervised action forecasting,” IEEE Transactions on Image Processing , vol. 29, pp. 8880–8891, 2020
2020
Later among the works it cites.
Y. Zhou, M. Wang, D. Liu, Z. Hu, and H. Zhang, “More grounded image captioning by distilling image-text matching model,” in CVPR , 2020
2020
Later among the works it cites.
B. Pan, H. Cai, D.-A. Huang, K.-H. Lee, A. Gaidon, E. Adeli, and J. C. Niebles, “Spatio-temporal graph for video captioning with knowledge distillation,” in CVPR , 2020
2020
Later among the works it cites.
Z. Zhang, Y. Shi, C. Yuan, B. Li, P. Wang, W. Hu, and Z. Zha, “Object relational graph with teacher-recommended learning for video captioning,” in CVPR , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
L. Zhou, H. Palangi, L. Zhang, H. Hu, J. Corso, and J. Gao, “Unified vision-language pre-training for image captioning and vqa,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, 2020, pp. 13 041–13 049
2020
Later among the works it cites.
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei et al. , “Oscar: Object-semantics aligned pre-training for vision-language tasks,” in European Conference on Computer Vision . Springer, 2020, pp. 121–137
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Yang, Y. Lu, J. Wang, X. Yin, D. Florencio, L. Wang, C. Zhang, L. Zhang, and J. Luo, “Tap: Text-aware pre-training for text-vqa and text-caption,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 8751–8761
2021
Later among the works it cites.