Fetching the paper…
Reading the bibliography…
Generative adversarial networks have led to significant advances in cross-modal/domain translation.
Signal estimation from modified short-time fourier transform
D. Griffin and J. Lim · 1984
Earlier work this paper cites.
Auditory-visual integration during multimodal object recognition in humans: a behavioral and electrophysiological study
M. H. Giard and F. Peronnet · 1999
Earlier work this paper cites.
The information bottleneck method
N. Tishby, F. C. Pereira, and W. Bialek · 2000
Earlier work this paper cites.
Hearing sounds, understanding actions: action representation in mirror neurons
E. Kohler, C. Keysers, M. A. Umilta, L. Fogassi, V. Gallese, and G. Rizzolatti · 2002
Earlier work this paper cites.
Beyond sensory images: Object-based representation in the human ventral pathway
P. Pietrini, M. L. Furey, E. Ricciardi, M. I. Gobbini, W.-H. C. Wu, L. Cohen, M. Guazzelli, and J. V. Haxby · 2004
Earlier work this paper cites.
Describing objects by their attributes
A. Farhadi, I. Endres, D. Hoiem, and D. Forsyth · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Y. LeCun and C. Cortes · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng · 2011
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 Dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Transient attributes for high-level understanding and editing of outdoor scenes
P.-Y. Laffont, Z. Ren, X. Tao, C. Qian, and J. Hays · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
M. Mirza and S. Osindero · 2014
Earlier work this paper cites.
Fine-grained visual comparisons with local learning
A. Yu and K. Grauman · 2014
Cited alongside, same era.
Microsoft COCO captions: Data collection and evaluation server
X. Chen, H. Fang, T. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Cited alongside, same era.
Librispeech: An ASR corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen, J. Chen, J. Chen, Z. Chen, M. Chrzanowski, A. Coates, G. Diamos, K. Ding, N. Du, E. Elsen, J. Engel, W. Fang, L. Fan, C. Fougner, L. Gao, C. Gong, A. Hannun, T. Han, L. V. Johannes, B. Jiang, C. Ju, B. Jun, P. LeGresley, L. Lin, J. Liu, Y. Liu, W. Li, X. Li, D. Ma, S. Narang, A. Ng, S. Ozair, Y. Peng, R. Prenger, S. Qian, Z. Quan, J. Raiman, V. Rao, S. Satheesh, D. Seetapun, S. Sengupta, K. Srinet, A. Sriram, H. Tang, L. Tang, C. Wang, J. Wang, K. Wang, Y. Wang, Z. Wang, Z. Wang, S. Wu, L. Wei, B. Xiao, W. Xie, Y. Xie, D. Yogatama, B. Yuan, J. Zhan, and Z. Zhu · 2016
Neural discrete representation learning
A. van den Oord, O. Vinyals, and k. kavukcuoglu · 2017
Later among the works it cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. Metaxas · 2017
Later among the works it cites.
Very deep convolutional networks for end-to-end speech recognition
Y. Zhang, W. Chan, and N. Jaitly · 2017
Later among the works it cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
J. Zhu, T. Park, P. Isola, and A. A. Efros · 2017
Later among the works it cites.
Toward multimodal image-to-image translation
J.-Y. Zhu, R. Zhang, D. Pathak, T. Darrell, A. A. Efros, O. Wang, and E. Shechtman · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generative adversarial text to image synthesis
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Generative visual manipulation on the natural image manifold
J.-Y. Zhu, P. Krähenbühl, E. Shechtman, and A. A. Efros · 2016
Cited alongside, same era.
Cvae-gan: Fine-grained image generation through asymmetric training
J. Bao, D. Chen, F. Wen, H. Li, and G. Hua · 2017
Cited alongside, same era.
One-sided unsupervised domain mapping
S. Benaim and L. Wolf · 2017
Cited alongside, same era.
Image-to-image translation with conditional adversarial networks
P. Isola, J. Zhu, T. Zhou, and A. A. Efros · 2017
Cited alongside, same era.
Conditional image synthesis with auxiliary classifier gans
A. Odena, C. Olah, and J. Shlens · 2017
Cited alongside, same era.
Later among the works it cites.
Stargan: Unified generative adversarial networks for multi-domain image-to-image translation
Y. Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo · 2018
Later among the works it cites.
Da-gan: Instance-level image translation by deep attention generative adversarial networks
S. Ma, J. Fu, C. Wen Chen, and T. Mei · 2018
Later among the works it cites.
Towards end-to-end prosody transfer for expressive speech synthesis with tacotron
R. J. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. J. Weiss, R. Clark, and R. A. Saurous · 2018
Later among the works it cites.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
Y. Wang, D. Stanton, Y. Zhang, R. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous · 2018
Later among the works it cites.
The microsoft 2017 conversational speech recognition system
W. Xiong, L. Wu, F. Alleva, J. Droppo, X. Huang, and A. Stolcke · 2018
Later among the works it cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Later among the works it cites.
A generative adversarial network for style modeling in a text-to-speech system
S. Ma, D. Mcduff, and Y. Song · 2019
Closest in time.