Fetching the paper…
Reading the bibliography…
We propose a novel class of language models, Latent Thought Models (LTMs), which incorporate explicit latent thought vectors that follow an explicit prior model in latent space.
The Language of Thought
Fodor, J. A · 1975
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Marcus, M., Santorini, B., and Marcinkiewicz, M. A · 1993
Earlier work this paper cites.
Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory
McClelland, J. L., McNaughton, B. L., and O’Reilly, R. C · 1995
Earlier work this paper cites.
An introduction to variational methods for graphical models
Jordan, M. I., Ghahramani, Z., Jaakkola, T. S., and Saul, L. K · 1999
Earlier work this paper cites.
The neural basis of lexicon and grammar in first and second language: The declarative/procedural model
Ullman, M. T · 2001
Earlier work this paper cites.
Contributions of memory circuits to language: The declarative/procedural model
Ullman, M. T · 2004
Earlier work this paper cites.
Bootstrapping in a language of thought: A formal model of numerical concept learning
Piantadosi, S. T., Tenenbaum, J. B., and Goodman, N. D · 2011
Earlier work this paper cites.
Machine Learning: A Probabilistic Perspective
Murphy, K. P · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T · 2013
Earlier work this paper cites.
Stochastic variational inference
Hoffman, M. D., Blei, D. M., Wang, C., and Paisley, J · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
A deep and tractable density estimator
Uria, B., Murray, I., and Larochelle, H · 2014
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
Using fast weights to attend to the recent past
Ba, J., Hinton, G. E., Mnih, V., Leibo, J. Z., and Ionescu, C · 2016
Earlier work this paper cites.
Generating sentences from a continuous space
Bowman, S. R., Vilnis, L., Vinyals, O., Dai, A. M., Jozefowicz, R., and Bengio, S · 2016
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Graves, A · 2016
Earlier work this paper cites.
What learning systems do intelligent agents need? complementary learning systems theory updated
Kumaran, D., Hassabis, D., and McClelland, J. L · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Variational inference: A review for statisticians
Blei, D. M., Kucukelbir, A., and McAuliffe, J. D · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Beam search strategies for neural machine translation
Freitag, M. and Al-Onaizan, Y · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
A discourse-aware attention model for abstractive summarization of long documents
Cohan, A., Dernoncourt, F., Kim, D. S., Bui, T., Kim, S., Chang, W., and Goharian, N · 2018
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., et al · 2022
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Later among the works it cites.
Continuous diffusion for categorical data
Dieleman, S., Sartran, L., Roshannai, A., Savinov, N., Ganin, Y., Richemond, P. H., Doucet, A., Strudel, R., Dyer, C., Durkan, C., et al · 2022
Later among the works it cites.
Han, X., Kumar, S., and Tsvetkov, Y · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dynamic evaluation of neural sequence models
Krause, B., Kahembwe, E., Murray, I., and Renals, S · 2018
Cited alongside, same era.
Spherical latent spaces for stable variational autoencoders
Xu, J. and Durrett, G · 2018
Cited alongside, same era.
Bayesian model-agnostic meta-learning
Yoon, J., Kim, T., Dia, O., Kim, S., Bengio, Y., and Ahn, S · 2018
Cited alongside, same era.
Openwebtext corpus
Gokaslan, A. and Cohen, V · 2019
Cited alongside, same era.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Don’t blame the elbo! a linear vae perspective on posterior collapse
Lucas, J., Tucker, G., Grosse, R., and Norouzi, M · 2019
Cited alongside, same era.
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
Autoregressive diffusion models
Hoogeboom, E., Gritsenko, A. A., Bastings, J., Poole, B., van den Berg, R., and Salimans, T · 2022
Later among the works it cites.
Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation
Kasai, J., Pappas, N., Peng, H., Cross, J., and Smith, N. A · 2022
Later among the works it cites.
Composing ensembles of pre-trained models via iterative consensus
Li, S., Du, Y., Tenenbaum, J. B., Torralba, A., and Mordatch, I · 2022
Later among the works it cites.
Latent diffusion energy-based model for interpretable text modeling
Yu, P., Xie, S., Ma, X., Jia, B., Pang, B., Gao, R., Zhu, Y., Zhu, S.-C., and Wu, Y. N · 2022
Later among the works it cites.
Star: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., Mu, J., and Goodman, N · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Later among the works it cites.
Amortizing intractable inference in large language models
Hu, E. J., Jain, M., Elmoznino, E., Kaddar, Y., Lajoie, G., Bengio, Y., and Malkin, N · 2023
Later among the works it cites.
Training chain-of-thought via latent-variable inference
Phan, D., Hoffman, M. D., Dohan, D., Douglas, S., Le, T. A., Parisi, A., Sountsov, P., Sutton, C., Vikram, S., and A Saurous, R · 2023
Later among the works it cites.
Diverse and faithful knowledge-grounded dialogue generation via sequential posterior inference
Xu, Y., Kong, D., Xu, D., Ji, Z., Pang, B., Fung, P., and Wu, Y. N · 2023
Later among the works it cites.
Training large language models to reason in a continuous latent space
Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J., and Tian, Y · 2024
Later among the works it cites.
Liger kernel: Efficient triton kernels for llm training
Hsu, P.-L., Dai, Y., Kothapalli, V., Song, Q., Tang, S., Zhu, S., Shimizu, S., Sahni, S., Ning, H., and Chen, Y · 2024
Later among the works it cites.
Discrete diffusion modeling by estimating the ratios of the data distribution
Lou, A., Meng, C., and Ermon, S · 2024
Later among the works it cites.
Simple and effective masked diffusion language models
Sahoo, S. S., Arriola, M., Schiff, Y., Gokaslan, A., Marroquin, E., Chiu, J. T., Rush, A., and Kuleshov, V · 2024
Later among the works it cites.
Simplified and generalized masked diffusion for discrete data
Shi, J., Han, K., Wang, Z., Doucet, A., and Titsias, M. K · 2024
Later among the works it cites.
Large concept models: Language modeling in a sentence representation space
The, L., Barrault, L., Duquenne, P.-A., Elbayad, M., Kozhevnikov, A., Alastruey, B., Andrews, P., Coria, M., Couairon, G., Costa-jussà, M. R., et al · 2024
Later among the works it cites.
Zheng, K., Chen, Y., Mao, H., Liu, M.-Y., Zhu, J., and Zhang, Q · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.