3d shapenets: A deep representation for volumetric shapes. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1912–1920
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 2015 · 1920
Earlier work this paper cites.
Stackgan++: Realistic image synthesis with stacked generative adversarial networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. 2018 · 1962
Earlier work this paper cites.
Speech recognition by machine: A review
D Raj Reddy. 1976 · 1976
Earlier work this paper cites.
Automatic generation of control signals for a parallel formant speech synthesizer. In ICASSP’76. IEEE International Conference on Acoustics, Speech, and Signal Processing , Vol. 1. IEEE, 690–693
P Seeviour, J Holmes, and M Judd. 1976 · 1976
Earlier work this paper cites.
Rule synthesis of speech from dyadic units. In ICASSP’77. IEEE International Conference on Acoustics, Speech, and Signal Processing , Vol. 2. IEEE, 568–570
Joseph Olive. 1977 · 1977
Earlier work this paper cites.
MITalk-79: The 1979 MIT text-to-speech system
Jonathan Allen, Sharon Hunnicutt, Rolf Carlson, and Bjorn Granstrom. 1979 · 1979
Earlier work this paper cites.
Correlation functions and computer simulations
Giorgio Parisi. 1981 · 1981
Earlier work this paper cites.
An introduction to the application of the theory of probabilistic functions of a Markov process to automatic speech recognition
Stephen E Levinson, Lawrence R Rabiner, and M Mohan Sondhi. 1983 · 1983
Earlier work this paper cites.
Learning internal representations by error propagation
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. 1985 · 1985
Earlier work this paper cites.
Machine translation: past, present, future
William John Hutchins. 1986 · 1986
Earlier work this paper cites.
Modular learning in neural networks.. In Aaai , Vol. 647. 279–284
Dana H Ballard. 1987 · 1987
Earlier work this paper cites.
An artificial intelligence approach to the conceptual description of videodisc images
Alan Philip Parkes. 1988 · 1988
Earlier work this paper cites.
The prototype CLORIS system: Describing, retrieving and discussing videodisc stills and sequences
Alan P Parkes. 1989b · 1989
Earlier work this paper cites.
Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones
Eric Moulines and Francis Charpentier. 1990 · 1990
Earlier work this paper cites.
Making the World Differentiable: On Using Self-Supervised Fully Recurrent N eu al Networks for Dynamic Reinforcement Learning and Planning in Non-Stationary Environm nts
Jiirgen Schmidhuber. 1990 · 1990
Earlier work this paper cites.
Hidden Markov models for speech recognition
Biing Hwang Juang and Laurence R Rabiner. 1991 · 1991
Earlier work this paper cites.
Minimal rules for articulatory speech synthesis
BJ Kröger. 1992 · 1992
Earlier work this paper cites.
Architecture of a multimedia information system for content-based retrieval. In Network and Operating System Support for Digital Audio and Video: Third International Workshop La Jolla, California, USA, November 12–13, 1992 Proceedings 3 . Springer, 387–392
Deborah Swanberg, Chiao Fe Shu, and Ramesh Jain. 1993 · 1992
Earlier work this paper cites.
Autoencoders, minimum description length and Helmholtz free energy
Geoffrey E Hinton and Richard Zemel. 1993 · 1993
Earlier work this paper cites.
Automatic structure visualization for video editing. In Proceedings of the INTERACT’93 and CHI’93 Conference on Human Factors in Computing Systems . 137–141
Hirotada Ueda, Takafumi Miyatake, Shigeo Sumino, and Akio Nagasaka. 1993 · 1993
Earlier work this paper cites.
Media streams: representing video for retrieval and repurposing. In Proceedings of the second ACM international conference on Multimedia . 478–479
Marc Davis. 1994 · 1994
Earlier work this paper cites.
Representations of knowledge in complex systems
Ulf Grenander and Michael I Miller. 1994 · 1994
Earlier work this paper cites.
IDIC: Assembling Video Sequences from Story Plans and Content Annotations.. In ICMCS . 30–36
Warren Sack and Marc Davis. 1994 · 1994
Earlier work this paper cites.
Structured video computing
Yoshinobu Tonomura, Akihito Akutsu, Yukinobu Taniguchi, and Gen Suzuki. 1994 · 1994
Earlier work this paper cites.
Conceptual indexing for video retrieval. In Working Notes of IJCAI Workshop on Intelligent Multimedia Information Retrieval, Montreal . 23–38
Andrew S Gordon and Eric A Domeshek. 1995 · 1995
Earlier work this paper cites.
AUTEUR: The application of video semantics and theme representation for automated film editing
Frank-Michael Nack. 1996 · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K Paliwal. 1997 · 1997
Earlier work this paper cites.
Collective dynamics of ‘small-world’networks
Duncan J Watts and Steven H Strogatz. 1998 · 1998
Earlier work this paper cites.
A general language model for information retrieval. In Proceedings of the eighth international conference on Information and knowledge management . 316–321
Fei Song and W Bruce Croft. 1999 · 1999
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent. 2000 · 2000
Earlier work this paper cites.
Fast texture synthesis using tree-structured vector quantization. In Proceedings of the 27th annual conference on Computer graphics and interactive techniques . 479–488
Li-Yi Wei and Marc Levoy. 2000 · 2000
Earlier work this paper cites.
Recurrent neural networks
Larry R Medsker and LC Jain. 2001 · 2001
Earlier work this paper cites.
Prospects for articulatory synthesis: A position paper. In 4th ISCA Tutorial and Research Workshop (ITRW) on Speech Synthesis
Christine H Shadle and Robert I Damper. 2001 · 2001
Earlier work this paper cites.
Statistical mechanics of complex networks
Réka Albert and Albert-László Barabási. 2002 · 2002
Earlier work this paper cites.
Woods. RE,(2002)“Digital Image Processing”
RC Gonzalez. 2006 · 2002
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Geoffrey E Hinton. 2002 · 2002
Earlier work this paper cites.
A new learning algorithm for mean field Boltzmann machines. In International Conference on Artificial Neural Networks . Springer, 351–357
Max Welling and Geoffrey E Hinton. 2002 · 2002
Earlier work this paper cites.
Dynamic textures
Gianfranco Doretto, Alessandro Chiuso, Ying Nian Wu, and Stefano Soatto. 2003 · 2003
Earlier work this paper cites.
Statistical phrase-based translation
Philipp Koehn, Franz J Och, and Daniel Marcu. 2003 · 2003
Earlier work this paper cites.
Towards automatic redeye effect removal
Bogdan Smolka, K Czubin, Jon Yngve Hardeberg, Kostas N Plataniotis, Marek Szczepanski, and Konrad Wojciechowski. 2003 · 2003
Earlier work this paper cites.
Auto cropping for digital photographs. In 2005 IEEE International Conference on Multimedia and Expo . IEEE, 4–pp
Mingju Zhang, Lei Zhang, Yanfeng Sun, Lin Feng, and Weiying Ma. 2005 · 2005
Earlier work this paper cites.
Mova Contour Moves “Motion Capture To Reality Capture”
Rick DeMott. 2006 · 2006
Earlier work this paper cites.
STRAIGHT, exploitation of the other aspect of VOCODER: Perceptually isomorphic decomposition of speech sounds
Hideki Kawahara. 2006 · 2006
Earlier work this paper cites.
Topic modeling: beyond bag-of-words. In Proceedings of the 23rd international conference on Machine learning . 977–984
Hanna M Wallach. 2006 · 2006
Earlier work this paper cites.
Automatic speech recognition and speech variability: A review
Mohamed Benzeghiba, Renato De Mori, Olivier Deroo, Stephane Dupont, Teodora Erbes, Denis Jouvet, Luciano Fissore, Pietro Laface, Alfred Mertins, Christophe Ris, et al · 2007
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation. In Proceedings of the 45th annual meeting of the association for computational linguistics companion volume proceedings of the demo and poster sessions . 177–180
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, et al · 2007
Earlier work this paper cites.
A survey of image classification methods and techniques for improving classification performance
Dengsheng Lu and Qihao Weng. 2007 · 2007
Earlier work this paper cites.
Natural image denoising with convolutional networks
Viren Jain and Sebastian Seung. 2008 · 2008
Earlier work this paper cites.
Learning hidden Markov models with hidden Markov trees as observation distributions
Diego H Milone and Leandro E Di Persia. 2008 · 2008
Earlier work this paper cites.
User generated content in the newsroom: Professional and organisational constraints on participatory journalism
Steve Paulussen and Pieter Ugille. 2008 · 2008
Earlier work this paper cites.
Training restricted Boltzmann machines using approximations to the likelihood gradient. In Proceedings of the 25th international conference on Machine learning . 1064–1071
Tijmen Tieleman. 2008 · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database. In CVPR
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
What is web 2.0
Tim O’reilly. 2009 · 2009
Earlier work this paper cites.
Building The Curious Faces Of ’Benjamin Button’
Laura Sydell. 2009 · 2009
Earlier work this paper cites.
Hybrid Hidden Markov Model and artificial neural network for automatic speech recognition. In 2009 Pacific-Asia Conference on Circuits, Communications and Systems . IEEE, 682–685
Xian Tang. 2009 · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In AISTATS
Michael Gutmann and Aapo Hyvärinen. 2010 · 2010
Earlier work this paper cites.
Handbook of natural language processing
Nitin Indurkhya and Fred J Damerau. 2010 · 2010
Earlier work this paper cites.
Kronecker graphs: an approach to modeling networks
Jure Leskovec, Deepayan Chakrabarti, Jon Kleinberg, Christos Faloutsos, and Zoubin Ghahramani. 2010 · 2010
Earlier work this paper cites.
Recurrent neural network based language model.. In Interspeech , Vol. 2. Makuhari, 1045–1048
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Two-phase kernel estimation for robust motion deblurring. In European conference on computer vision . Springer, 157–170
Li Xu and Jiaya Jia. 2010 · 2010
Earlier work this paper cites.
Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition
George E Dahl, Dong Yu, Li Deng, and Alex Acero. 2011 · 2011
Earlier work this paper cites.
Apertium: a free/open-source platform for rule-based machine translation
Mikel L Forcada, Mireia Ginestí-Rosell, Jacob Nordfalk, Jim O’Regan, Sergio Ortiz-Rojas, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Gema Ramírez-Sánchez, and Francis M Tyers. 2011 · 2011
Earlier work this paper cites.
The neural autoregressive distribution estimator. In Proceedings of the fourteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 29–37
Hugo Larochelle and Iain Murray. 2011 · 2011
Earlier work this paper cites.
Data-driven response generation in social media. In Empirical Methods in Natural Language Processing (EMNLP)
Alan Ritter, Colin Cherry, and Bill Dolan. 2011 · 2011
Earlier work this paper cites.
A hierarchical, context-dependent neural network architecture for improved phone recognition. In 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5040–5043
László Tóth. 2011 · 2011
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Pascal Vincent. 2011 · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks. In NeurIPS
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012 · 2012
Earlier work this paper cites.
Multimodal learning with deep boltzmann machines
Nitish Srivastava and Russ R Salakhutdinov. 2012 · 2012
Earlier work this paper cites.
Image denoising and inpainting with deep neural networks
Junyuan Xie, Linli Xu, and Enhong Chen. 2012 · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Original
Alex Graves. 2013 · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks. In 2013 IEEE international conference on acoustics, speech and signal processing . Ieee, 6645–6649
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. 2013 · 2013
Earlier work this paper cites.
Recurrent continuous translation models. In Proceedings of the 2013 conference on empirical methods in natural language processing . 1700–1709
Nal Kalchbrenner and Phil Blunsom. 2013 · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Original
Diederik P Kingma and Max Welling. 2013 · 2013
Earlier work this paper cites.
Social Book Search: The Impact of Professional and User-Generated Content on Book Suggestions.. In DIR . 38–39
Marijn Koolen, Jaap Kamps, Gabriella Kazai, et al · 2013
Earlier work this paper cites.
Learning out-of-vocabulary words in automatic speech recognition
Long Qin. 2013 · 2013
Earlier work this paper cites.
RNADE: The real-valued neural autoregressive density-estimator
Benigno Uria, Iain Murray, and Hugo Larochelle. 2013 · 2013
Earlier work this paper cites.
Furious 7 (2015)
WIKIPEDIA. 2013 · 2013
Earlier work this paper cites.
Statistical parametric speech synthesis using deep neural networks. In 2013 ieee international conference on acoustics, speech and signal processing . IEEE, 7962–7966
Heiga Ze, Andrew Senior, and Mike Schuster. 2013 · 2013
Earlier work this paper cites.
Gradient-based Wiener filter for image denoising
Xiaobo Zhang, Xiangchu Feng, Weiwei Wang, Shunli Zhang, and Qunfeng Dong. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Original
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Original
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Original
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Learning a deep convolutional network for image super-resolution. In European conference on computer vision . Springer, 184–199
Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. 2014 · 2014
Earlier work this paper cites.
TTS synthesis with bidirectional LSTM based recurrent neural networks. In Fifteenth annual conference of the international speech communication association
Yuchen Fan, Yao Qian, Feng-Long Xie, and Frank K Soong. 2014 · 2014
Earlier work this paper cites.
Generative adversarial nets. In NeurIPS
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Original
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille. 2014 · 2014
Earlier work this paper cites.
The First News Report on the L.A. Earthquake Was Written by a Robot
Will Oremus. 2014 · 2014
Earlier work this paper cites.
On the training aspects of deep neural network (DNN) for parametric TTS synthesis. In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3829–3833
Yao Qian, Yuchen Fan, Wenping Hu, and Frank K Soong. 2014 · 2014
Earlier work this paper cites.
Image denoising using new adaptive based median filters
Original
Suman Shrestha. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Deep convolutional neural network for image deconvolution
Li Xu, Jimmy S Ren, Ce Liu, and Jiaya Jia. 2014 · 2014
Earlier work this paper cites.
Neural shape codes for 3D model retrieval
Song Bai, Xiang Bai, Wenyu Liu, and Fabio Roli. 2015 · 2015
Earlier work this paper cites.
Deep colorization. In Proceedings of the IEEE international conference on computer vision . 415–423
Zezhou Cheng, Qingxiong Yang, and Bin Sheng. 2015 · 2015
Earlier work this paper cites.
Nice: Non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2625–2634
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015 · 2015
Earlier work this paper cites.
Image super-resolution using deep convolutional networks
Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. 2015 · 2015
Earlier work this paper cites.
Learning deep sigmoid belief networks with data augmentation. In Artificial Intelligence and Statistics . PMLR, 268–276
Zhe Gan, Ricardo Henao, David Carlson, and Lawrence Carin. 2015 · 2015
Earlier work this paper cites.
A neural algorithm of artistic style
Original
Leon A Gatys, Alexander S Ecker, and Matthias Bethge. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3128–3137
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Applied regression: An introduction . Vol. 22
Colin Lewis-Beck and Michael Lewis-Beck. 2015 · 2015
Earlier work this paper cites.
Stacked denoising autoencoder and dropout together to prevent overfitting in deep neural network. In 2015 8th international congress on image and signal processing (CISP) . IEEE, 697–701
Jianglin Liang and Ruifang Liu. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Original
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Original
Alec Radford, Luke Metz, and Soumith Chintala. 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition. In ICLR
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
Learning a convolutional neural network for non-uniform motion blur removal. In Proceedings of the IEEE conference on computer vision and pattern recognition . 769–777
Jian Sun, Wenfei Cao, Zongben Xu, and Jean Ponce. 2015 · 2015
Earlier work this paper cites.
Going deeper with convolutions. In CVPR
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3156–3164
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Earlier work this paper cites.
Deep convolutional architecture for natural image denoising. In 2015 International Conference on Wireless Communications & Signal Processing (WCSP) . IEEE, 1–4
Xuejiao Wang, Qiuyan Tao, Lianghao Wang, Dongxiao Li, and Ming Zhang. 2015 · 2015
Earlier work this paper cites.
Acoustic modeling in statistical parametric speech synthesis-from HMM to LSTM-RNN
Heiga Zen. 2015 · 2015
Earlier work this paper cites.
Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4470–4474
Heiga Zen and Haşim Sak. 2015 · 2015
Earlier work this paper cites.
Tutorial - What is a Variational Autoencoder?
Jaan Altosaar. 2016 · 2016
Earlier work this paper cites.
Unitary evolution recurrent neural networks. In International conference on machine learning . PMLR, 1120–1128
Martin Arjovsky, Amar Shah, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Layer normalization
Original
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
3d-r2n2: A unified approach for single and multi-view 3d object reconstruction. In European conference on computer vision . Springer, 628–644
Christopher B Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 2016 · 2016
Earlier work this paper cites.
Wav2letter: an end-to-end convnet-based speech recognition system
Original
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve. 2016 · 2016
Earlier work this paper cites.
Image style transfer using convolutional neural networks. In CVPR
Leon A Gatys, Alexander S Ecker, and Matthias Bethge. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In CVPR
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
How YouTube developed into a successful platform for user-generated content
Margaret Holland. 2016 · 2016
Earlier work this paper cites.
Can active memory replace attention?
Łukasz Kaiser and Samy Bengio. 2016 · 2016
Earlier work this paper cites.
Neural machine translation in linear time
Original
Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
A review on automatic speech recognition architecture and approaches
S Karpagavalli and Edy Chandra. 2016 · 2016
Earlier work this paper cites.
Power of consumers using social media: Examining the influences of brand-related user-generated content on Facebook
Angella J Kim and Kim KP Johnson. 2016 · 2016
Earlier work this paper cites.
Convolutional network for attribute-driven and identity-preserving human face generation
Original
Mu Li, Wangmeng Zuo, and David Zhang. 2016a · 2016
Earlier work this paper cites.
Deep identity-aware transfer of facial attributes
Original
Mu Li, Wangmeng Zuo, and David Zhang. 2016b · 2016
Earlier work this paper cites.
Generating images from captions with attention
Elman Mansimov, Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov. 2016 · 2016
Earlier work this paper cites.
WORLD: a vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa. 2016 · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles. In ECCV
Mehdi Noroozi and Paolo Favaro. 2016 · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
Original
Ankur P Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Earlier work this paper cites.
Automatic Speech Recognition and its Applications
Pratiksha C Raut and Seema U Deoghare. 2016 · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis. In International conference on machine learning . PMLR, 1060–1069
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016 · 2016
Earlier work this paper cites.
Defining Web 3.0: opportunities and challenges
Riaan Rudman and Rikus Bruwer. 2016 · 2016
Earlier work this paper cites.
Native language influence during second language acquisition: A large-scale learner corpus analysis. In Proceedings of the Pacific Second Language Research Forum (PacSLRF 2016) . 175–188
Itamar Shatz. 2017 · 2016
Earlier work this paper cites.
Neural autoregressive distribution estimation
Benigno Uria, Marc-Alexandre Côté, Karol Gregor, Iain Murray, and Hugo Larochelle. 2016 · 2016
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio. In The 9th ISCA Speech Synthesis Workshop
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al · 2016
Earlier work this paper cites.
Fully automatic image colorization based on Convolutional Neural Network. In 2016 23rd International Conference on Pattern Recognition (ICPR) . IEEE, 3691–3696
Domonkos Varga and Tamás Szirányi. 2016 · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Original
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
Image deblurring with blur kernel estimation in RGB channels. In 2016 IEEE International Conference on Digital Signal Processing (DSP) . IEEE, 681–684
Xianqiu Xu, Hongqing Liu, Yong Li, and Yi Zhou. 2016 · 2016
Earlier work this paper cites.
Attribute2image: Conditional image generation from visual attributes. In European conference on computer vision . Springer, 776–791
Xinchen Yan, Jimei Yang, Kihyuk Sohn, and Honglak Lee. 2016 · 2016
Earlier work this paper cites.
Generative visual manipulation on the natural image manifold. In European conference on computer vision . Springer, 597–613
Jun-Yan Zhu, Philipp Krähenbühl, Eli Shechtman, and Alexei A Efros. 2016 · 2016
Earlier work this paper cites.
Multi-source neural translation
Original
Barret Zoph and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Image distortion detection using convolutional neural network. In 2017 4th IAPR Asian Conference on Pattern Recognition (ACPR) . IEEE, 220–225
Namhyuk Ahn, Byungkon Kang, and Kyung-Ah Sohn. 2017 · 2017
Earlier work this paper cites.
Subtitles for the Deaf or Hard-of-Hearing (SDH) - Subtitles, Closed Captions, and SDH
ai media. 2017 · 2017
Earlier work this paper cites.
Massive Exploration of Neural Machine Translation Architectures
Original
Denny Britz, Anna Goldie, Thang Luong, and Quoc Le. 2017 · 2017
Earlier work this paper cites.
You said that?
Original
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Improved speech reconstruction from silent video. In Proceedings of the IEEE International Conference on Computer Vision Workshops . 455–462
Ariel Ephrat, Tavi Halperin, and Shmuel Peleg. 2017 · 2017
Earlier work this paper cites.
Vid2speech: speech reconstruction from silent video. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5095–5099
Ariel Ephrat and Shmuel Peleg. 2017 · 2017
Earlier work this paper cites.
Convolutional sequence to sequence learning. In International conference on machine learning . PMLR, 1243–1252
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin. 2017 · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
Densely connected convolutional networks. In CVPR
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q. Weinberger. 2017 · 2017
Earlier work this paper cites.
How Google is making music with arti4cial intelligence
Matthew Hutson. 2017 · 2017
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, et al · 2017
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen. 2017 · 2017
Earlier work this paper cites.
Newspaper companies’ determinants in adopting robot journalism
Daewon Kim and Seongcheol Kim. 2017 · 2017
Earlier work this paper cites.
Unsupervised visual attribute transfer with reconfigurable generative adversarial networks
Original
Taeksoo Kim, Byoungjip Kim, Moonsu Cha, and Jiwon Kim. 2017 · 2017
Earlier work this paper cites.
Obamanet: Photo-realistic lip-sync from text
Original
Rithesh Kumar, Jose Sotelo, Kundan Kumar, Alexandre de Brébisson, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
Photo-realistic single image super-resolution using a generative adversarial network. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4681–4690
Christian Ledig, Lucas Theis, Ferenc Huszár, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al · 2017
Earlier work this paper cites.
Learning a model of facial shape and expression from 4D scans
Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero. 2017 · 2017
Earlier work this paper cites.
Not all dialogues are created equal: Instance weighting for neural conversational models
Original
Pierre Lison and Serge Bibauw. 2017 · 2017
Earlier work this paper cites.
An investigation of brand-related user-generated content on Twitter
Xia Liu, Alvin C Burns, and Yingjian Hou. 2017 · 2017
Earlier work this paper cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning. In Proceedings of the IEEE conference on computer vision and pattern recognition . 375–383
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Attentive semantic video generation using captions. In Proceedings of the IEEE international conference on computer vision . 1426–1434
Tanya Marwah, Gaurav Mittal, and Vineeth N Balasubramanian. 2017 · 2017
Earlier work this paper cites.
Sync-draw: Automatic video generation using deep recurrent attentive architectures. In Proceedings of the 25th ACM international conference on Multimedia . 1096–1104
Gaurav Mittal, Tanya Marwah, and Vineeth N Balasubramanian. 2017 · 2017
Earlier work this paper cites.
To create what you tell: Generating videos from captions. In Proceedings of the 25th ACM international conference on Multimedia . 1789–1798
Yingwei Pan, Zhaofan Qiu, Ting Yao, Houqiang Li, and Tao Mei. 2017 · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. 2017b · 2017
Earlier work this paper cites.
Temporal generative adversarial nets with singular value clipping. In Proceedings of the IEEE international conference on computer vision . 2830–2839
Masaki Saito, Eiichi Matsumoto, and Shunta Saito. 2017 · 2017
Earlier work this paper cites.
Pixelcnn++: Improving the pixelcnn with discretized logistic mixture likelihood and other modifications
Original
Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P Kingma. 2017 · 2017
Earlier work this paper cites.
A deep reinforcement learning chatbot
Original
Iulian V Serban, Chinnadhurai Sankar, Mathieu Germain, Saizheng Zhang, Zhouhan Lin, Sandeep Subramanian, Taesup Kim, Michael Pieper, Sarath Chandar, Nan Rosemary Ke, et al · 2017
Earlier work this paper cites.
Learning residual images for face attribute manipulation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4030–4038
Wei Shen and Rujie Liu. 2017 · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman. 2017 · 2017
Earlier work this paper cites.