D. H. Ackley, G. E. Hinton, and T. J. Sejnowski, “A Learning Algorithm for Boltzmann Machines,” Cognitive Science , vol. 9, no. 1, pp. 147–169, 1985
1985
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning Internal Representations by Error Propagation,” Neurocomputing , 1987
1987
Earlier work this paper cites.
R. J. Williams and D. Zipser, “A Learning Algorithm for Continually Running Fully Recurrent Neural Networks,” Neural Computation , vol. 1, no. 2, pp. 270–280, 1989
1989
Earlier work this paper cites.
Y. Bengio, P. Simard, and P. Frasconi, “Learning Long-Term Dependencies with Gradient Descent is Difficult,” IEEE trans. neural netw. , vol. 5, no. 2, pp. 157–166, 1994
1994
Earlier work this paper cites.
T. Landauer and S. Dumais, “A Solution to Plato’s Problem: The Latent Semantic Analysis Theory of Acquisition, Induction, and Representation of Knowledge.” Psychological Review , vol. 104, no. 2, p. 211, 1997
1997
Earlier work this paper cites.
S. Hochreiter, “The Vanishing Gradient Problem During Learning Recurrent Neural Nets and Problem Solutions,” Int. J. Uncertain. Fuzziness Knowledge-Based Syst. , vol. 6, no. 02, pp. 107–116, 1998
1998
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: a Method for Automatic Evaluation of Machine Translation,” in Proc. of ACL , 2002
2002
Earlier work this paper cites.
G. Doddington, “Automatic Evaluation of Machine Translation Quality Using N-gram Co-Occurrence Statistics,” in Proc. of ICHLT , 2002
2002
Earlier work this paper cites.
C.-Y. Lin, “ROUGE: A Package for Automatic Evaluation of Summaries,” in Proc. of ACL , 2004
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,” in Proc. of ACL Workshop , 2005
2005
Earlier work this paper cites.
F. Pedregosa et al., “Scikit-learn: Machine Learning in Python,” JMLR , vol. 12, pp. 2825–2830, 2011
2011
Earlier work this paper cites.
“GridSearchCV in ScikitLearn,” https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GridSearchCV.html , 2011
2011
Earlier work this paper cites.
“Logistic Regression Classifier in ScikitLearn,” https://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegression.html , 2011
2011
Earlier work this paper cites.
A. Graves, “Long Short-Term Memory,” Springer , pp. 37–45, 2012
2012
Earlier work this paper cites.
V. Rus and M. Lintean, “An Optimal Assessment of Natural Language Student Input Using Word-to-Word Similarity Metrics,” in Proc. of ITS , 2012
2012
Earlier work this paper cites.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient Estimation of Word Representations in Vector Space,” in Proc. of ICLR Workshop , 2013
2013
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “GloVe: Global Vectors for Word Representation,” in Proc. of EMNLP , 2014
2014
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in Proc. of ICLR , 2015
2015
Earlier work this paper cites.
M. J. Kusner, Y. Sun, N. I. Kolkin, and K. Q. Weinberger, “From Word Embeddings to Document Distances,” in Proc. of ICML , 2015
2015
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural Machine Translation of Rare Words with Subword Units,” in Proc. of ACL , 2016
2016
Earlier work this paper cites.
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “SQuAD: 100,000+ Questions for Machine Comprehension of Text,” in Proc. of EMNLP , 2016
2016
Earlier work this paper cites.
“Hugging Face: The AI community building the future,” https://huggingface.co/models , 2016
2016
Earlier work this paper cites.
N. Mrkšić, D. O. Séaghdha, B. Thomson, M. Gašić, L. Rojas-Barahona, P.-H. Su, D. Vandyke, T.-H. Wen, and S. Young, “Counter-fitting Word Vectors to Linguistic Constraints,” in Proc. of NAACL-HLT , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in Proc. of NeurIPS , 2017
2017
Earlier work this paper cites.
L. Leppänen, M. Munezero, M. Granroth-Wilding, and H. Toivonen, “Data-Driven News Generation for Automated Journalism,” in Proc. of INLG , 2017
2017
Earlier work this paper cites.
Y. Yao, B. Viswanath, J. Cryan, H. Zheng, and B. Y. Zhao, “Automated Crowdturfing Attacks and Defenses in Online Review Systems,” in Proc. of CCS , 2017
2017
Earlier work this paper cites.
S. Wiseman, S. M. Shieber, and A. M. Rush, “Challenges in Data-to-Document Generation,” in Proc. of EMNLP , 2017
2017
Earlier work this paper cites.
G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy, “RACE: Large-scale ReAding Comprehension Dataset From Examinations,” in Proc. of EMNLP , 2017
2017
Earlier work this paper cites.