Fetching the paper…
Reading the bibliography…
We study the probabilistic modeling performed by Autoregressive Large Language Models (LLMs) through the angle of time directionality, addressing a question first raised in (Shannon, 1951).
Insertion-based Decoding with automatically Inferred Generation Order, October 2019a
Gu, J., Liu, Q., and Cho, K · 1902
Earlier work this paper cites.
Insertion Transformer: Flexible Sequence Generation via Insertion Operations, February 2019
Stern, M., Chan, W., Kiros, J., and Uszkoreit, J · 1902
Earlier work this paper cites.
Non-Monotonic Sequential Text Generation, October 2019
Welleck, S., Brantley, K., Daumé III, H., and Cho, K · 1902
Earlier work this paper cites.
Levenshtein Transformer, October 2019b
Gu, J., Wang, C., and Zhao, J · 1905
Earlier work this paper cites.
Mansimov, E., Wang, A., Welleck, S., and Cho, K · 1905
Earlier work this paper cites.
KERMIT: Generative Insertion-Based Modeling for Sequences, June 2019a
Chan, W., Kitaev, N., Guu, K., Stern, M., and Uszkoreit, J · 1906
Earlier work this paper cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding, January 2020
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., and Le, Q. V · 1906
Earlier work this paper cites.
An Empirical Study of Generation Order for Machine Translation, October 2019b
Chan, W., Stern, M., Kiros, J., and Uszkoreit, J · 1910
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, September 2023
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 1910
Earlier work this paper cites.
Unsupervised Cross-lingual Representation Learning at Scale, April 2020
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V · 1911
Earlier work this paper cites.
Sequence Modeling with Unconstrained Generation Order, October 2019
Emelianenko, D., Voita, E., and Serdyukov, P · 1911
Earlier work this paper cites.
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data, November 2019
Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzmán, F., Joulin, A., and Grave, E · 1911
Earlier work this paper cites.
Prediction and entropy of printed english
Shannon, C. E · 1951
Earlier work this paper cites.
Elicitation of Personal Probabilities and Expectations
Savage, L. J · 1971
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2001
Earlier work this paper cites.
Confidence scoring based on backward language models
Duchateau, J., Demuynck, K., and Wambacq, P · 2002
Earlier work this paper cites.
A Theory of Usable Information Under Computational Constraints, February 2020
Xu, Y., Zhao, S., Song, J., Stewart, R., and Ermon, S · 2002
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Strictly Proper Scoring Rules, Prediction, and Estimation
Gneiting, T. and Raftery, A. E · 2007
Earlier work this paper cites.
LOGARITHMIC MARKETS CORING RULES FOR MODULAR COMBINATORIAL INFORMATION AGGREGATION
Hanson, R · 2012
Cited alongside, same era.
Sequence to Sequence Learning with Neural Networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Cited alongside, same era.
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Agreement on Target-bidirectional Neural Machine Translation
Liu, L., Utiyama, M., Finch, A., and Sumita, E · 2016
Cited alongside, same era.
Backward and Forward Language Modeling for Constrained Sentence Generation, January 2016
Mou, L., Yan, R., Li, G., Zhang, L., and Jin, Z · 2016
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, May 2019
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Later among the works it cites.
Decoupled Weight Decay Regularization, January 2019
Loshchilov, I. and Hutter, F · 2019
Later among the works it cites.
LSTM vs. GRU vs. Bidirectional RNN for script generation
Mangal, S., Joshi, P., and Modak, R · 2019
Later among the works it cites.
Elements of Causal Inference: Foundations and Learning Algorithms
Peters, J., Janzing, D., and Schölkopf, B · 2019
Later among the works it cites.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural Machine Translation of Rare Words with Subword Units, June 2016
Sennrich, R., Haddow, B., and Birch, A · 2016
Cited alongside, same era.
Order Matters: Sequence to sequence for sets, February 2016
Vinyals, O., Bengio, S., and Kudlur, M · 2016
Cited alongside, same era.
Direct Methods for Sparse Matrices
Duff, I. S., Erisman, A. M., and Reid, J. K · 2017
Cited alongside, same era.
SGDR: Stochastic Gradient Descent with Warm Restarts, May 2017
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Parallel WaveNet: Fast High-Fidelity Speech Synthesis, November 2017
Oord, A. v. d., Li, Y., Babuschkin, I., Simonyan, K., Vinyals, O., Kavukcuoglu, K., Driessche, G. v. d., Lockhart, E., Cobo, L. C., Stimberg, F., Casagrande, N., Grewe, D., Noury, S., Dieleman, S., Elsen, E., Kalchbrenner, N., Zen, H., Graves, A., King, H., Walters, T., Belov, D., and Hassabis, D · 2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
The Importance of Generation Order in Language Modeling
Ford, N., Duckworth, D., Norouzi, M., and Dahl, G · 2018
Cited alongside, same era.
Karpathy, A · 2020
Later among the works it cites.
Towards Causal Representation Learning, February 2021
Scholkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., and Bengio, Y · 2021
Later among the works it cites.
Training Compute-Optimal Large Language Models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., Driessche, G. v. d., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Later among the works it cites.
Step-unrolled Denoising Autoencoders for Text Generation, April 2022
Savinov, N., Chung, J., Binkowski, M., Elsen, E., and Oord, A. v. d · 2022
Later among the works it cites.
Emergent Abilities of Large Language Models, June 2022
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W · 2022
Later among the works it cites.
Boolformer: Symbolic Regression of Logic Functions with Transformers, September 2023
d’Ascoli, S., Bengio, S., Susskind, J., and Abbé, E · 2023
Later among the works it cites.
Bayesian Flow Networks, August 2023
Graves, A., Srivastava, R. K., Atkinson, T., and Gomez, F · 2023
Later among the works it cites.
Teaching Arithmetic to Small Transformers, July 2023
Lee, N., Sreenivasan, K., Lee, J. D., Lee, K., and Papailiopoulos, D · 2023
Later among the works it cites.
Meet in the Middle: A New Pre-training Paradigm, March 2023
Nguyen, A., Karampatziakis, N., and Chen, W · 2023
Later among the works it cites.
GPT-4 Technical Report, March 2023
OpenAI · 2023
Later among the works it cites.
Eliciting Language Model Behaviors using Reverse Language Models
Pfau, J., Infanger, A., Sheshadri, A., Panda, A., Huebner, C., and Michael, J · 2023
Later among the works it cites.
Are Emergent Abilities of Large Language Models a Mirage?, May 2023
Schaeffer, R., Miranda, B., and Koyejo, S · 2023
Later among the works it cites.
Positional Description Matters for Transformers Arithmetic, November 2023
Shen, R., Bubeck, S., Eldan, R., Lee, Y. T., Li, Y., and Zhang, Y · 2023
Later among the works it cites.
This week I trained an 800K transformer to learn 5 digit multiplication., September 2023
Thomas Ahle · 2023
Later among the works it cites.
A Survey on Non-Autoregressive Generation for Neural Machine Translation and Beyond, July 2023
Xiao, Y., Wu, L., Guo, J., Li, J., Zhang, M., Qin, T., and Liu, T.-y · 2023
Later among the works it cites.