Fetching the paper…
Reading the bibliography…
The video action segmentation task is regularly explored under weaker forms of supervision, such as transcript supervision, where a list of actions is easier to obtain than dense frame-wise labels.
V. I. Levenshtein
1966
Earlier work this paper cites.
R. Hadsell, S. Chopra, and Y. LeCun, “Dimensionality reduction by learning an invariant mapping,” in
2006
Earlier work this paper cites.
G. Chechik, V. Sharma, U. Shalit, and S. Bengio, “Large scale online learning of image similarity through ranking.”
2010
Earlier work this paper cites.
J. Deng, A. C. Berg, and L. Fei-Fei, “Hierarchical semantic indexing for large scale image retrieval,” in
2011
Earlier work this paper cites.
B. Siddiquie, R. S. Feris, and L. S. Davis, “Image ranking and retrieval based on multi-attribute queries,” in
2011
Earlier work this paper cites.
S. Stein and S. J. McKenna, “Combining embedded accelerometers with computer vision for recognizing food preparation activities,” in
2013
Earlier work this paper cites.
P. Bojanowski, R. Lajugie, F. Bach, I. Laptev, J. Ponce, C. Schmid, and J. Sivic, “Weakly supervised action labeling in videos under ordering constraints,” in
2014
Earlier work this paper cites.
H. Kuehne, A. Arslan, and T. Serre, “The language of actions: Recovering the syntax and semantics of goal-directed human activities,” in
2014
Earlier work this paper cites.
K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder–decoder for statistical machine translation,” in
2014
Earlier work this paper cites.
J. Wang, Y. Song, T. Leung, C. Rosenberg, J. Wang, J. Philbin, B. Chen, and Y. Wu, “Learning fine-grained image similarity with deep ranking,” in
2014
Earlier work this paper cites.
C. Xing, D. Wang, C. Liu, and Y. Lin, “Normalized word embedding and orthogonal transform for bilingual word translation,” in
2015
Earlier work this paper cites.
D.-A. Huang, L. Fei-Fei, and J. C. Niebles, “Connectionist temporal modeling for weakly supervised action labeling,” in
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
A. Richard, H. Kuehne, and J. Gall, “Weakly supervised action learning with rnn based fine-to-coarse modeling,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in
2017
Earlier work this paper cites.
A. Richard, H. Kuehne, A. Iqbal, and J. Gall, “Neuralnetwork-viterbi: A framework for weakly supervised video learning,” in
2018
Cited alongside, same era.
L. Ding and C. Xu, “Weakly-supervised action segmentation with iterative soft boundary assignment,” in
2018
Cited alongside, same era.
H. Coskun, D. J. Tan, S. Conjeti, N. Navab, and F. Tombari, “Human motion analysis with deep metric learning,” in
2018
Cited alongside, same era.
F. Sener and A. Yao, “Unsupervised learning and segmentation of complex activities from video,” in
2018
Cited alongside, same era.
J. Li, P. Lei, and S. Todorovic, “Weakly supervised energy-based learning for action segmentation,” in
2019
Cited alongside, same era.
M. S. Hutchinson and V. N. Gadepally, “Video action understanding: A tutorial,”
2021
Later among the works it cites.
X. Chang, F. Tung, and G. Mori, “Learning discriminative prototypes with dynamic time warping,” in
2021
Later among the works it cites.
Y. Souri, M. Fayyaz, L. Minciullo, G. Francesca, and J. Gall, “Fast weakly supervised action segmentation using mutual consistency,”
2021
Later among the works it cites.
F. Yi, H. Wen, and T. Jiang, “Asformer: Transformer for action segmentation,” in
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,”
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C.-Y. Chang, D.-A. Huang, Y. Sui, L. Fei-Fei, and J. C. Niebles, “D3tw: Discriminative differentiable dynamic time warping for weakly supervised action alignment and segmentation,” in
2019
Cited alongside, same era.
2019
Cited alongside, same era.
N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” in
2019
Cited alongside, same era.
B. Barz and J. Denzler, “Hierarchy-based image embeddings for semantic image retrieval,” in
2019
Cited alongside, same era.
A. Kukleva, H. Kuehne, F. Sener, and J. Gall, “Unsupervised learning of action classes with continuous temporal embedding,” in
2019
Cited alongside, same era.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Zhang, X. Li, C. Liu, B. Shuai, Y. Zhu, B. Brattoli, H. Chen, I. Marsic, and J. Tighe, “Vidtr: Video transformer without convolutions,” in
2021
Later among the works it cites.
D. Neimark, O. Bar, M. Zohar, and D. Asselmann, “Video transformer network,”
2021
Later among the works it cites.
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. Torr
2021
Later among the works it cites.
R. Girdhar and K. Grauman, “Anticipative video transformer,”
2021
Later among the works it cites.
2021
Later among the works it cites.
R. Tan, H. Xu, K. Saenko, and B. A. Plummer, “Logan: Latent graph co-attention network for weakly-supervised video moment retrieval,” in
2092
Closest in time.