Fetching the paper…
Reading the bibliography…
As language models become more powerful, training and evaluation are increasingly bottlenecked by the data and metrics used for a particular task.
A technique for the measurement of attitudes
R. Likert · 1932
Earlier work this paper cites.
Speech understanding systems: A summary of results of the five-year research effort. department of computer science, 1977
D. R. Reddy et al · 1977
Earlier work this paper cites.
Optimum polynomial retrieval functions based on the probability ranking principle
N. Fuhr · 1989
Earlier work this paper cites.
Automatic combination of multiple ranked retrieval systems
B. T. Bartell, G. W. Cottrell, and R. K. Belew · 1994
Earlier work this paper cites.
Optimizing search engines using clickthrough data
T. Joachims · 2002
Earlier work this paper cites.
Hedge trimmer: A parse-and-trim approach to headline generation
B. Dorr, D. Zajic, and R. Schwartz · 2003
Earlier work this paper cites.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
C.-Y. Lin and F. J. Och · 2004
Earlier work this paper cites.
Accurately interpreting clickthrough data as implicit feedback
T. Joachims, L. Granka, B. Pan, H. Hembrooke, and G. Gay · 2005
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Learning to rank for information retrieval
T.-Y. Liu · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Semi-supervised sequence learning
A. M. Dai and Q. V. Le · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2015
Earlier work this paper cites.
A neural attention model for abstractive sentence summarization
A. M. Rush, S. Chopra, and J. Weston · 2015
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
D. Bahdanau, P. Brakel, K. Xu, A. Goyal, R. Lowe, J. Pineau, A. Courville, and Y. Bengio · 2016
Earlier work this paper cites.
Abstractive sentence summarization with attentive recurrent neural networks
S. Chopra, M. Auli, and A. M. Rush · 2016
Earlier work this paper cites.
Deep neural networks for youtube recommendations
P. Covington, J. Adams, and E. Sargin · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Teaching machines to describe images with natural language feedback
S. Fidler et al · 2017
Earlier work this paper cites.
Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
N. Jaques, S. Gu, D. Bahdanau, J. M. Hernández-Lobato, R. E. Turner, and D. Eck · 2017
Earlier work this paper cites.
Tuning recurrent neural networks with reinforcement learning
N. Jaques, S. Gu, R. E. Turner, and D. Eck · 2017
Earlier work this paper cites.
Reinforcement learning for bandit neural machine translation with simulated human feedback
K. Nguyen, H. Daumé III, and J. Boyd-Graber · 2017
Cited alongside, same era.
A deep reinforced model for abstractive summarization
R. Paulus, C. Xiong, and R. Socher · 2017
Cited alongside, same era.
The limits of automatic summarisation according to rouge
N. Schluter · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Get to the point: Summarization with pointer-generator networks
A. See, P. J. Liu, and C. D. Manning · 2017
Learning from dialogue after deployment: Feed yourself, chatbot!
B. Hancock, A. Bordes, P.-E. Mazare, and J. Weston · 2019
Later among the works it cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2019
Later among the works it cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2019
Later among the works it cites.
Neural text summarization: A critical evaluation
W. Kryscinski, N. S. Keskar, B. McCann, C. Xiong, and R. Socher · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Tl; dr: Mining reddit to learn automatic summarization
M. Völske, M. Potthast, S. Syed, and B. Stein · 2017
Cited alongside, same era.
The price of debiasing automatic metrics in natural language evaluation
A. T. Chaganty, S. Mussman, and P. Liang · 2018
Cited alongside, same era.
Towards coherent and cohesive long-form text generation
W. S. Cho, P. Zhang, Y. Zhang, X. Li, M. Galley, C. Brockett, M. Wang, and J. Gao · 2018
Cited alongside, same era.
Supervising strong learners by amplifying weak experts
P. Christiano, B. Shlegeris, and D. Amodei · 2018
Cited alongside, same era.
Banditsum: Extractive summarization as a contextual bandit
Y. Dong, Y. Shen, E. Crawford, H. van Hoof, and J. C. K. Cheung · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
B. Ibarz, J. Leike, T. Pohlen, G. Irving, S. Legg, and D. Amodei · 2018
Cited alongside, same era.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer · 2019
Later among the works it cites.
Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons
M. Li, J. Weston, and S. Roller · 2019
Later among the works it cites.
Finding generalizable evidence by learning to convince q&a models
E. Perez, S. Karamcheti, R. Fergus, J. Weston, D. Kiela, and K. Cho · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Later among the works it cites.
Generalization in generation: A closer look at exposure bias
F. Schmidt · 2019
Later among the works it cites.
Mass: Masked sequence to sequence pre-training for language generation
K. Song, X. Tan, T. Qin, J. Lu, and T.-Y. Liu · 2019
Later among the works it cites.
Neural text generation with unlikelihood training
S. Welleck, I. Kulikov, S. Roller, E. Dinan, K. Cho, and J. Weston · 2019
Later among the works it cites.
S. Yi, R. Goel, C. Khatri, A. Cervone, T. Chung, B. Hedayatnia, A. Venkatesh, R. Gabriel, and D. Hakkani-Tur · 2019
Later among the works it cites.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
J. Zhang, Y. Zhao, M. Saleh, and P. J. Liu · 2019
Later among the works it cites.
Abstract text summarization with a convolutional seq2seq model
Y. Zhang, D. Li, Y. Wang, Y. Fang, and W. Xiao · 2019
Later among the works it cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Later among the works it cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Closest in time.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
J. Dodge, G. Ilharco, R. Schwartz, A. Farhadi, H. Hajishirzi, and N. Smith · 2020
Closest in time.
Reward-rational (implicit) choice: A unifying formalism for reward learning
H. J. Jeon, S. Milli, and A. D. Dragan · 2020
Closest in time.
On faithfulness and factuality in abstractive summarization, 2020
J. Maynez, S. Narayan, B. Bohnet, and R. McDonald · 2020
Closest in time.
Leveraging pre-trained checkpoints for sequence generation tasks
S. Rothe, S. Narayan, and A. Severyn · 2020
Closest in time.
Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training
Y. Yan, W. Qi, Y. Gong, D. Liu, N. Duan, J. Chen, R. Zhang, and M. Zhou · 2020
Closest in time.
Trading off diversity and quality in natural language generation
H. Zhang, D. Duckworth, D. Ippolito, and A. Neelakantan · 2020
Closest in time.
W. Zhou and K. Xu · 2020
Closest in time.