Fetching the paper…
Reading the bibliography…
In recent years, natural language processing (NLP) has got great development with deep learning techniques.
So, D. R., Liang, C., & Le, Q. V. (2019). The evolved transformer. arXiv preprint arXiv:1901.11117
1901
Earlier work this paper cites.
1901
Earlier work this paper cites.
So, D. R., Liang, C., & Le, Q. V. (2019). The evolved transformer. arXiv preprint arXiv:1901.11117
1901
Earlier work this paper cites.
1902
Earlier work this paper cites.
1902
Earlier work this paper cites.
1903
Earlier work this paper cites.
1904
Earlier work this paper cites.
1904
Earlier work this paper cites.
1905
Earlier work this paper cites.
1906
Earlier work this paper cites.
1906
Earlier work this paper cites.
1906
Earlier work this paper cites.
1909
Earlier work this paper cites.
JORDAN, M. (1986). Serial Order; a parallel distributed processing approach. ICS Report 8604, UC San Diego
1986
Earlier work this paper cites.
Allen, R. B. (1987, June). Several studies on natural language and back-propagation. In Proceedings of the IEEE First International Conference on Neural Networks (Vol. 2, No. S 335, p. 341). IEEE Piscataway, NJ
1987
Earlier work this paper cites.
Pollack, J. B. (1990). Recursive distributed representations. Artificial Intelligence, 46(1-2), 77-105
1990
Earlier work this paper cites.
Elman, J. L. (1990). Finding structure in time. Cognitive science, 14(2), 179-211
1990
Earlier work this paper cites.
Chrisman, L. (1991). Learning recursive distributed representations for holistic computation. Connection Science, 3(4), 345-366
1991
Earlier work this paper cites.
Bengio, Y., Simard, P., & Frasconi, P. (1994). Learning long-term dependencies with gradient descent is difficult. IEEE transactions on neural networks, 5(2), 157-166
1994
Earlier work this paper cites.
Gage, P. (1994). A new algorithm for data compression. The C Users Journal, 12(2), 23-38
1994
Earlier work this paper cites.
Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural computation, 9(8), 1735-1780
1997
Earlier work this paper cites.
Miller, G. (1998). WordNet: An electronic lexical database. MIT press
1998
Earlier work this paper cites.
Och, F. J., Tillmann, C., & Ney, H. (1999). Improved alignment models for statistical machine translation. In 1999 Joint SIGDAT Conference on Empirical Methods in Natural Language Processing and Very Large Corpora
1999
Earlier work this paper cites.
Song, F., & Croft, W. B. (1999, November). A general language model for information retrieval. In Proceedings of the eighth international conference on Information and knowledge management (pp. 316-321). ACM
1999
Earlier work this paper cites.
Rosenfeld, R. (2000). Two decades of statistical language modeling: Where do we go from here?. Proceedings of the IEEE, 88(8), 1270-1278
2000
Earlier work this paper cites.
Stolcke, A. (2002). SRILM-an extensible language modeling toolkit. In Seventh international conference on spoken language processing
2002
Earlier work this paper cites.
Koehn, P., Och, F. J., & Marcu, D. (2003, May). Statistical phrase-based translation. In Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology-Volume 1 (pp. 48-54). Association for Computational Linguistics
2003
Earlier work this paper cites.
Bengio, Y., Ducharme, R., Vincent, P., & Jauvin, C. (2003). A neural probabilistic language model. Journal of machine learning research, 3(Feb), 1137-1155
2003
Earlier work this paper cites.
Bengio, Y., & Senécal, J. S. (2003, January). Quick Training of Probabilistic Neural Nets by Importance Sampling. In AISTATS(pp. 1-9)
2003
Earlier work this paper cites.
Koehn, P. (2004, September). Pharaoh: a beam search decoder for phrase-based statistical machine translation models. In Conference of the Association for Machine Translation in the Americas (pp. 115-124). Springer, Berlin, Heidelberg
2004
Earlier work this paper cites.
Chiang, D. (2005, June). A hierarchical phrase-based model for statistical machine translation. In Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics (pp. 263-270). Association for Computational Linguistics
2005
Earlier work this paper cites.
Schwenk, H., & Gauvain, J. L. (2005, October). Training neural network language models on very large corpora. In Proceedings of the conference on Human Language Technology and Empirical Methods in Natural Language Processing (pp. 201-208). Association for Computational Linguistics
2005
Earlier work this paper cites.
Morin, F., & Bengio, Y. (2005, January). Hierarchical probabilistic neural network language model. In Aistats (Vol. 5, pp. 246-252)
2005
Earlier work this paper cites.
Schwenk, H., Dchelotte, D., & Gauvain, J. L. (2006, July). Continuous space language models for statistical machine translation. In Proceedings of the COLING/ACL on Main conference poster sessions (pp. 723-730). Association for Computational Linguistics
2006
Earlier work this paper cites.
Teh, Y. W. (2006, July). A hierarchical Bayesian language model based on Pitman-Yor processes. In Proceedings of the 21st International Conference on Computational Linguistics and the 44th annual meeting of the Association for Computational Linguistics (pp. 985-992). Association for Computational Linguistics
2006
Earlier work this paper cites.
Wallach, H. M. (2006, June). Topic modeling: beyond bag-of-words. In Proceedings of the 23rd international conference on Machine learning (pp. 977-984). ACM
2006
Earlier work this paper cites.
Koehn, P., Hoang, H., Birch, A., Callison-Burch, C., Federico, M., Bertoldi, N., … & Dyer, C. (2007, June). Moses: Open source toolkit for statistical machine translation. In Proceedings of the 45th annual meeting of the association for computational linguistics companion volume proceedings of the demo and poster sessions (pp. 177-180)
2007
Earlier work this paper cites.
Schwenk, H. (2007). Continuous space language models. Computer Speech & Language, 21(3), 492-518
2007
Earlier work this paper cites.
2007
Earlier work this paper cites.
Galley, M., and Manning, C. D. (2008, October). A simple and effective hierarchical phrase reordering model. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (pp. 848-856). Association for Computational Linguistics
2008
Earlier work this paper cites.
Bengio, Y., & Senécal, J. S. (2008). Adaptive importance sampling to accelerate training of a neural probabilistic language model. IEEE Transactions on Neural Networks, 19(4), 713-722
2008
Earlier work this paper cites.
Federico, M., Bertoldi, N., & Cettolo, M. (2008). IRSTLM: an open source toolkit for handling large scale language models. In Ninth Annual Conference of the International Speech Communication Association
2008
Earlier work this paper cites.
Chiang, D., Knight, K., & Wang, W. (2009, May). 11,001 new features for statistical machine translation. In Proceedings of human language technologies: The 2009 annual conference of the north american chapter of the association for computational linguistics (pp. 218-226). Association for Computational Linguistics
2009
Earlier work this paper cites.
Mnih, A., & Hinton, G. E. (2009). A scalable hierarchical distributed language model. In Advances in neural information processing systems (pp. 1081-1088)
2009
Earlier work this paper cites.
Mikolov, T., Karafiát, M., Burget, L., Černocký, J., & Khudanpur, S. (2010). Recurrent neural network based language model. In Eleventh annual conference of the international speech communication association
2010
Earlier work this paper cites.
Forcada, M. L., Ginestí-Rosell, M., Nordfalk, J., O’Regan, J., Ortiz-Rojas, S., Pérez-Ortiz, J. A., … & Tyers, F. M. (2011). Apertium: a free/open-source platform for rule-based machine translation. Machine translation, 25(2), 127-144
2011
Earlier work this paper cites.
Heafield, K. (2011, July). KenLM: Faster and smaller language model queries. In Proceedings of the sixth workshop on statistical machine translation (pp. 187-197). Association for Computational Linguistics
2011
Earlier work this paper cites.
Collobert, R., Weston, J., Bottou, L., Karlen, M., Kavukcuoglu, K., & Kuksa, P. (2011). Natural language processing (almost) from scratch. Journal of machine learning research, 12(Aug), 2493-2537
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
Schuster, M., & Nakajima, K. (2012, March). Japanese and korean voice search. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 5149-5152). IEEE
2012
Earlier work this paper cites.
Schwenk, H. (2012, December). Continuous space translation models for phrase-based statistical machine translation. In Proceedings of COLING 2012: Posters (pp. 1071-1080)
2012
Earlier work this paper cites.
Son, L. H., Allauzen, A., & Yvon, F. (2012, June). Continuous space translation models with neural networks. In Proceedings of the 2012 conference of the north american chapter of the association for computational linguistics: Human language technologies (pp. 39-48). Association for Computational Linguistics
2012
Cited alongside, same era.
Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems (pp. 1097-1105)
2012
Cited alongside, same era.
Kalchbrenner, N., & Blunsom, P. (2013). Recurrent continuous translation models. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (pp. 1700-1709)
2013
Cited alongside, same era.
2013
Cited alongside, same era.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mikolov, T., Yih, W. T., & Zweig, G. (2013). Linguistic regularities in continuous space word representations. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 746-751)
2013
Cited alongside, same era.
Green, S., Wang, S., Cer, D., & Manning, C. D. (2013). Fast and adaptive online training of feature-rich translation models. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (Vol. 1, pp. 311-321)
2013
Cited alongside, same era.
Boulanger-Lewandowski, N., Bengio, Y., & Vincent, P. (2013, November). Audio Chord Recognition with Recurrent Neural Networks. In ISMIR (pp. 335-340)
2013
Cited alongside, same era.
2013
Cited alongside, same era.
Auli, M., Galley, M., Quirk, C.,& Zweig, G. (2013). Joint language and translation modeling with recurrent neural networks
2013
Cited alongside, same era.
Vaswani, A., Zhao, Y., Fossum, V., & Chiang, D. (2013). Decoding with large-scale neural language models improves translation. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (pp. 1387-1392)
2013
Cited alongside, same era.
Maas, A. L., Hannun, A. Y., & Ng, A. Y. (2013, June). Rectifier nonlinearities improve neural network acoustic models. In Proc. icml (Vol. 30, No. 1, p. 3)
2013
Cited alongside, same era.
2014
Cited alongside, same era.
2016
Later among the works it cites.
Duong, L., Anastasopoulos, A., Chiang, D., Bird, S., & Cohn, T. (2016, June). An attentional model for speech translation without transcription. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 949-959)
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
He, D., Xia, Y., Qin, T., Wang, L., Yu, N., Liu, T. Y., & Ma, W. Y. (2016). Dual learning for machine translation. In Advances in Neural Information Processing Systems (pp. 820-828)
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778)
2016
Later among the works it cites.
2016
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., … & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems(pp. 5998-6008)
2017
Later among the works it cites.
2017
Later among the works it cites.
Nakazawa, T., Higashiyama, S., Ding, C., Mino, H., Goto, I., Kazawa, H., … & Kurohashi, S. (2017, November). Overview of the 4th Workshop on Asian Translation. In Proceedings of the 4th Workshop on Asian Translation (WAT2017) (pp. 1-54)
2017
Later among the works it cites.
Choi, H., Cho, K., & Bengio, Y. (2017). Context-dependent word representation for neural machine translation. Computer Speech & Language, 45, 149-160
2017
Later among the works it cites.
Lee, J., Cho, K., & Hofmann, T. (2017). Fully character-level neural machine translation without explicit segmentation. Transactions of the Association for Computational Linguistics, 5, 365-378
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Gehring, J., Auli, M., Grangier, D., Yarats, D., & Dauphin, Y. N. (2017, August). Convolutional sequence to sequence learning. In Proceedings of the 34th International Conference on Machine Learning-Volume 70 (pp. 1243-1252). JMLR. org
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Dauphin, Y. N., Fan, A., Auli, M., & Grangier, D. (2017, August). Language modeling with gated convolutional networks. In Proceedings of the 34th International Conference on Machine Learning-Volume 70 (pp. 933-941). JMLR. org
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Luong, M. T. (2017).Neural Machine Translation. Unpublished doctoral dissertation, Stanford University, Stanford, CA 94305
2017
Later among the works it cites.
2018
Later among the works it cites.
Zhao, Y., Zhang, J., He, Z., Zong, C., & Wu, H. (2018). Addressing Troublesome Words in Neural Machine Translation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (pp. 391-400)
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
Domhan, T. (2018, July). How much attention do you need? a granular analysis of neural machine translation architectures. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1799-1808)
2018
Later among the works it cites.
Young, T., Hazarika, D., Poria, S., & Cambria, E. (2018). Recent trends in deep learning based natural language processing. ieee Computational intelligenCe magazine, 13(3), 55-75
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. URL https://s3-us-west-2. amazonaws. com/openai-assets/researchcovers/languageunsupervised/language understanding paper. pdf
2018
Later among the works it cites.
2018
Later among the works it cites.
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., … & Bengio, Y. (2015, June). Show, attend and tell: Neural image caption generation with visual attention. In International conference on machine learning (pp. 2048-2057)
2057
Closest in time.
Chung, J., Gulcehre, C., Cho, K., & Bengio, Y. (2015, June). Gated feedback recurrent neural networks. In International Conference on Machine Learning (pp. 2067-2075)
2075
Closest in time.