Fetching the paper…
Reading the bibliography…
Measuring the quality of a generated sequence against a set of references is a central problem in many learning frameworks, be it to compute a score, to assign a reward, or to perform discrimination.
What does bert look at? an analysis of bert’s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D · 1906
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Williams, R. J. and Zipser, D · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Re-evaluating the role of Bleu in machine translation research
Callison-Burch, C., Osborne, M., and Koehn, P · 2006
Earlier work this paper cites.
K-means++: The advantages of careful seeding
Arthur, D. and Vassilvitskii, S · 2007
Earlier work this paper cites.
Book review: Natural language processing with python by steven bird, ewan Klein, and edward loper
Elhadad, M · 2010
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Cho, K., van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Earlier work this paper cites.
Kiros, R., Zhu, Y., Salakhutdinov, R., Zemel, R. S., Torralba, A., Urtasun, R., and Fidler, S · 2015
Earlier work this paper cites.
From word embeddings to document distances
Kusner, M. J., Sun, Y., Kolkin, N. I., and Weinberger, K. Q · 2015
Earlier work this paper cites.
Truly exploring multiple references for machine translation evaluation
Qin, Y. and Specia, L · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Ranzato, M., Chopra, S., Auli, M., and Zaremba, W · 2015
Earlier work this paper cites.
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A · 2016
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
Bahdanau, D., Brakel, P., Xu, K., Goyal, A., Lowe, R., Pineau, J., Courville, A. C., and Bengio, Y · 2016
Earlier work this paper cites.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Cited alongside, same era.
Professor forcing: A new algorithm for training recurrent networks
Goyal, A., Lamb, A., Zhang, Y., Zhang, S., Courville, A. C., and Bengio, Y · 2016
Cited alongside, same era.
Self-critical sequence training for image captioning
Rennie, S. J., Marcheret, E., Mroueh, Y., Ross, J., and Goel, V · 2016
Cited alongside, same era.
Building end-to-end dialogue systems using generative hierarchical neural network models
Serban, I. V., Sordoni, A., Bengio, Y., Courville, A. C., and Pineau, J · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., Klingner, J., Shah, A., Johnson, M., Liu, X., Kaiser, L., Gouws, S., Kato, Y., Kudo, T., Kazawa, H., Stevens, K., Kurian, G., Patil, N., Wang, W., Young, C., Smith, J., Riesa, J., Rudnick, A., Vinyals, O., Corrado, G., Hughes, M., and Dean, J · 2016
Visualizing and measuring the geometry of BERT
Coenen, A., Reif, E., Yuan, A., Kim, B., Pearce, A., Viégas, F. B., and Wattenberg, M · 2019
Later among the works it cites.
How contextual are contextualized word representations? comparing the geometry of BERT, ELMo, and GPT-2 embeddings
Ethayarajh, K · 2019
Later among the works it cites.
Discofuse: A large-scale dataset for discourse-based sentence fusion
Geva, M., Malmi, E., Szpektor, I., and Berant, J · 2019
Later among the works it cites.
A structural probe for finding syntax in word representations
Hewitt, J. and Manning, C. D · 2019
Later among the works it cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Forbes, M., and Choi, Y · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Seqgan: Sequence generative adversarial nets with policy gradient
Yu, L., Zhang, W., Wang, J., and Yu, Y · 2017
Cited alongside, same era.
Looking for elmo’s friends: Sentence-level pretraining beyond language modeling
Bowman, S. R., Pavlick, E., Grave, E., Durme, B. V., Wang, A., Hula, J., Xia, P., Pappagari, R., McCoy, R. T., Patel, R., Kim, N., Tenney, I., Huang, Y., Yu, K., Jin, S., and Chen, B · 2018
Cited alongside, same era.
Caccia, M., Caccia, L., Fedus, W., Larochelle, H., Pineau, J., and Charlin, L · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y. N · 2018
Cited alongside, same era.
Maskgan: Better text generation via filling in the ______
Fedus, W., Goodfellow, I. J., and Dai, A. M · 2018
Cited alongside, same era.
Learning to write with cooperative discriminators
Holtzman, A., Buys, J., Forbes, M., Bosselut, A., Golub, D., and Choi, Y · 2018
Cited alongside, same era.
Later among the works it cites.
Neural text summarization: A critical evaluation
Kryscinski, W., Keskar, N. S., McCann, B., Xiong, C., and Socher, R · 2019
Later among the works it cites.
Fully unsupervised crosslingual semantic textual similarity metric based on BERT for identifying parallel data
Lo, C.-k. and Simard, M · 2019
Later among the works it cites.
Latent normalizing flows for discrete sequences
M. Ziegler, Z. and M. Rush, A · 2019
Later among the works it cites.
Putting evaluation in context: Contextual embeddings improve machine translation evaluation
Mathur, N., Baldwin, T., and Cohn, T · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Generalization in generation: A closer look at exposure bias
Schmidt, F · 2019
Later among the works it cites.
BERT has a mouth, and it must speak: BERT as a markov random field language model
Wang, A. and Cho, K · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Later among the works it cites.
Beyond BLEU:training neural machine translation with semantic similarity
Wieting, J., Berg-Kirkpatrick, T., Gimpel, K., and Neubig, G · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 2019
Later among the works it cites.
Bertscore: Evaluating text generation with BERT
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2019
Later among the works it cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Zhao, W., Peyrard, M., Liu, F., Gao, Y., Meyer, C. M., and Eger, S · 2019
Later among the works it cites.