Fetching the paper…
Reading the bibliography…
With the development of media and networking technologies, multimedia applications ranging from feature presentation in a cinema setting to video on demand to interactive video conferencing are in great demand.
US Patent 5,430,485
D. E. Lankford and M. S. Deiss, “Audio/video synchronization in a digital transmission system,” July 4 1995 · 1995
Earlier work this paper cites.
US Patent 6,122,668
P. Y. Teng, B. A. Thompson, and F. A. Tobagi, “Synchronization of audio and video signals in a live multicast in a lan,” Sept. 19 2000 · 2000
Earlier work this paper cites.
J. Weston, S. Chopra, and A. Bordes, “Memory networks,” 2014
2014
Earlier work this paper cites.
E. Marcheret, G. Potamianos, J. Vopicka, and V. Goel, “Detecting audio-visual synchrony using deep neural networks,” in
2015
Earlier work this paper cites.
S. Sukhbaatar, J. Weston, R. Fergus,
2015
Earlier work this paper cites.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in
2015
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in
2015
Earlier work this paper cites.
2015
Cited alongside, same era.
S. Sharma, R. Kiros, and R. Salakhutdinov, “Action recognition using visual attention,”
2015
Cited alongside, same era.
M. Noroozi and P. Favaro, “Unsupervised learning of visual representations by solving jigsaw puzzles,” in
2016
Cited alongside, same era.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in
2016
Cited alongside, same era.
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola, “Stacked attention networks for image question answering,” in
2016
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Later among the works it cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in
2017
Later among the works it cites.
2018
Closest in time.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,”
2018
Closest in time.
J. Zang, L. Wang, Z. Liu, Q. Zhang, G. Hua, and N. Zheng, “Attention-based temporal weighted convolutional neural network for action recognition,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
J. S. Chung, A. W. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild.,” in
2017
Cited alongside, same era.
A. Mazaheri, D. Zhang, and M. Shah, “Video fill in the blank using lr/rl lstms with spatial-temporal attentions,”
Cited in the paper.
2018
Closest in time.