Fetching the paper…
Reading the bibliography…
Cross-modality generation is an emerging topic that aims to synthesize data in one modality based on information in a different modality.
Relations between two sets of variates
Hotelling, H.: · 1936
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D.E., Hinton, G.E., Williams, R.J., et al.: · 1988
Earlier work this paper cites.
Phoneme recognition using time-delay neural networks
Waibel, A.H., Hanazawa, T., Hinton, G.E., Shikano, K., Lang, K.J.: · 1989
Earlier work this paper cites.
Look who’s talking: Speaker detection using video and audio correlation
Cutler, R., Davis, L.S.: · 2000
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P.: · 2004
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
Cooke, M., Barker, J., Cunningham, S., Shao, X.: · 2006
Earlier work this paper cites.
Audiovisual database of spoken american english
Richie, S., Warburton, C., Carter, M.: · 2009
Earlier work this paper cites.
The natural statistics of audiovisual speech
Chandrasekaran, C., Trubanova, A., Stillittano, S., Caplier, A., Ghazanfar, A.A.: · 2009
Earlier work this paper cites.
Dlib-ml: A machine learning toolkit
King, D.E.: · 2009
Earlier work this paper cites.
A new approach to cross-modal multimedia retrieval
Rasiwasia, N., Pereira, J.C., Coviello, E., Doyle, G., Lanckriet, G.R.G., Levy, R., Vasconcelos, N.: · 2010
Earlier work this paper cites.
Baby talk: Understanding and generating simple image descriptions
Kulkarni, G., Premraj, V., Dhar, S., Li, S., Choi, Y., Berg, A.C., Berg, T.L.: · 2011
Earlier work this paper cites.
A no-reference image blur metric based on the cumulative probability of blur detection (CPBD)
Narvekar, N.D., Karam, L.J.: · 2011
Cited alongside, same era.
A thousand frames in just a few words: lingual description of videos through latent topics and sparse object stitching
Das, P., Xu, C., Doell, R., Corso, J.J.: · 2013
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I.J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A.C., Bengio, Y.: · 2014
Cited alongside, same era.
Conditional generative adversarial nets
Mirza, M., Osindero, S.: · 2014
Cited alongside, same era.
Photo-real talking head with deep bidirectional LSTM
Fan, B., Wang, L., Soong, F.K., Xie, L.: · 2015
Cited alongside, same era.
Lip reading in the wild
Chung, J.S., Zisserman, A.: · 2016
Later among the works it cites.
Generating videos with scene dynamics
Vondrick, C., Pirsiavash, H., Torralba, A.: · 2016
Later among the works it cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., Abbeel, P.: · 2016
Later among the works it cites.
Perceptual losses for real-time style transfer and super-resolution
Johnson, J., Alahi, A., Fei-Fei, L.: · 2016
Later among the works it cites.
Deep cross-modal audio-visual generation
Chen, L., Srivastava, S., Duan, Z., Xu, C.: · 2017
Later among the works it cites.
Synthesizing obama: learning lip sync from audio
Suwajanakorn, S., Seitz, S.M., Kemelmacher-Shlizerman, I.: · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Garrido, P., Valgaerts, L., Sarmadi, H., Steiner, I., Varanasi, K., Pérez, P., Theobalt, C.: · 2015
Cited alongside, same era.
Flownet: Learning optical flow with convolutional networks
Dosovitskiy, A., Fischer, P., Ilg, E., Häusser, P., Hazirbas, C., Golkov, V., van der Smagt, P., Cremers, D., Brox, T.: · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., Sun, J.: · 2015
Cited alongside, same era.
Visually indicated sounds
Owens, A., Isola, P., McDermott, J., Torralba, A., Adelson, E.H., Freeman, W.T.: · 2016
Cited alongside, same era.
Generative adversarial text to image synthesis
Reed, S.E., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., Lee, H.: · 2016
Cited alongside, same era.
Virtual immortality: Reanimating characters from TV shows
Charles, J., Magee, D.R., Hogg, D.C.: · 2016
Cited alongside, same era.
You said that?
Chung, J.S., Jamaludin, A., Zisserman, A.: · 2017
Later among the works it cites.
Lipnet: End-to-end sentence-level lipreading
Assael, Y.M., Shillingford, B., Whiteson, S., de Freitas, N.: · 2017
Later among the works it cites.
Conditional image synthesis with auxiliary classifier gans
Odena, A., Olah, C., Shlens, J.: · 2017
Later among the works it cites.
Lip reading sentences in the wild
Son Chung, J., Senior, A., Vinyals, O., Zisserman, A.: · 2017
Later among the works it cites.