Fetching the paper…
Reading the bibliography…
Autoregressive models (ARMs) have become the workhorse for sequence generation tasks, since many problems can be modeled as next-token prediction.
The analysis of permutations
Plackett, R. L · 1975
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Stein’s method for discrete Gibbs measures
Eichelsbacher, P. and Reinert, G · 2008
Earlier work this paper cites.
Salimans, T. and Knowles, D. A · 2014
Earlier work this paper cites.
A deep and tractable density estimator
Uria, B., Murray, I., and Larochelle, H · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Learning deep generative models of graphs
Li, Y., Vinyals, O., Dyer, C., Pascanu, R., and Battaglia, P · 2018
Earlier work this paper cites.
Fréchet chemnet distance: A metric for generative models for molecules in drug discovery
Preuer, K., Renz, P., Unterthiner, T., Hochreiter, S., and Klambauer, G · 2018
Earlier work this paper cites.
Graphrnn: Generating realistic graphs with deep auto-regressive models
You, J., Ying, R., Ren, X., Hamilton, W., and Leskovec, J · 2018
Earlier work this paper cites.
How powerful are graph neural networks?
Xu, K., Hu, W., Leskovec, J., and Jegelka, S · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Structured denoising diffusion models in discrete state-spaces
Austin, J., Johnson, D. D., Ho, J., Tarlow, D., and Van Den Berg, R · 2021
Cited alongside, same era.
Order matters: Probabilistic modeling of node sequence for graph generation
Chen, X., Han, X., Hu, J., Ruiz, F., and Liu, L · 2021
Cited alongside, same era.
Benchmarking graph neural networks
Dwivedi, V. P., Joshi, C. K., Luu, A. T., Laurent, T., Bengio, Y., and Bresson, X · 2023
Later among the works it cites.
Autoregressive diffusion model for graph generation
Kong, L., Cui, J., Sun, H., Zhuang, Y., Prakash, B. A., and Zhang, C · 2023
Later among the works it cites.
Molecule generation using transformers and policy gradient reinforcement learning
Mazuz, E., Shtar, G., Shapira, B., and Rokach, L · 2023
Later among the works it cites.
Digress: Discrete denoising diffusion for graph generation
Vignac, C., Krawczuk, I., Siraudin, A., Wang, B., Cevher, V., and Frossard, P · 2023
Later among the works it cites.
Variational flow matching for graph generation
Eijkelboom, F., Bartosh, G., Naesseth, C. A., Welling, M., and van de Meent, J.-W · 2024
Later among the works it cites.
Discrete diffusion modeling by estimating the ratios of the data distribution
Lou, A., Meng, C., and Ermon, S · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Autoregressive diffusion models
Hoogeboom, E., Gritsenko, A. A., Bastings, J., Poole, B., van den Berg, R., and Salimans, T · 2022
Cited alongside, same era.
Score-based generative modeling of graphs via the system of stochastic differential equations
Jo, J., Lee, S., and Hwang, S. J · 2022
Cited alongside, same era.
Gradient estimation with discrete Stein operators
Shi, J., Zhou, Y., Hwang, J., Titsias, K. M., and Mackey, L · 2022
Cited alongside, same era.
Let there be order: Rethinking ordering in autoregressive graph generation
Bu, J., Mehrab, K. S., and Karpatne, A · 2023
Cited alongside, same era.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y
Cited in the paper.
Buy 4 REINFORCE samples, get a baseline for free!
Kool, W., Hoof, H. V., and Welling, M
Cited in the paper.
Stochastic Beams and Where To Find Them: The Gumbel-Top-k Trick for Sampling Sequences Without Replacement
Kool, W., van Hoof, H., and Welling, M
Cited in the paper.
Later among the works it cites.
Your absorbing discrete diffusion secretly models the conditional distributions of clean data
Ou, J., Nie, S., Xue, K., Zhu, F., Sun, J., Li, Z., and Li, C · 2024
Later among the works it cites.
Simple and effective masked diffusion language models
Sahoo, S. S., Arriola, M., Schiff, Y., Gokaslan, A., Marroquin, E., Chiu, J. T., Rush, A., and Kuleshov, V · 2024
Later among the works it cites.
Simplified and generalized masked diffusion for discrete data
Shi, J., Han, K., Wang, Z., Doucet, A., and Titsias, M. K · 2024
Later among the works it cites.