Highly accurate protein structure prediction with AlphaFold
Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Zídek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., Back, T., Petersen, S., Reiman, D., Clancy, E., Zielinski, M., Steinegger, M., Pacholska, M., Berghammer, T., Bodenstein, S., Silver, D., Vinyals, O., Senior, A. W., Kavukcuoglu, K., Kohli, P., and Hassabis, D · 2021
Later among the works it cites.
On density estimation with diffusion models
Kingma, D. P., Salimans, T., Poole, B., and Ho, J · 2021
Later among the works it cites.
Generative spoken language modeling from raw audio
Original
Lakhotia, K., Kharitonov, E., Hsu, W.-N., Adi, Y., Polyak, A., Bolte, B., Nguyen, T.-A., Copet, J., Baevski, A., Mohamed, A., and Dupoux, E · 2021
Later among the works it cites.
LUNA: Linear unified nested attention
Ma, X., Kong, X., Wang, S., Zhou, C., May, J., Ma, H., and Zettlemoyer, L · 2021
Later among the works it cites.
Generating images with sparse representations
Nash, C., Menick, J., Dieleman, S., and Battaglia, P. W · 2021
Later among the works it cites.
Hierarchical Transformers are more efficient language models
Original
Nawrot, P., Tworkowski, S., Tyrolski, M., Kaiser, L., Wu, Y., Szegedy, C., and Michalewski, H · 2021
Later among the works it cites.
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N., and Kong, L · 2021
Later among the works it cites.
Speech resynthesis from discrete disentangled self-supervised representations
Original
Polyak, A., Adi, Y., Copet, J., Kharitonov, E., Lakhotia, K., Hsu, W.-N., Mohamed, A., and Dupoux, E · 2021
Later among the works it cites.
Shortformer: Better language modeling using shorter inputs
Press, O., Smith, N. A., and Lewis, M · 2021
Later among the works it cites.
Self-attention does not need O ( n 2 ) O(n^{2}) memory
Original
Rabe, M. N. and Staats, C · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training Gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P.-S., Glaese, A., Welbl, J., Dathathri, S., Huang, S., Uesato, J., Mellor, J., Higgins, I., Creswell, A., McAleese, N., Wu, A., Elsen, E., Jayakumar, S., Buchatskaya, E., Budden, D., Sutherland, E., Simonyan, K., Paganini, M., Sifre, L., Martens, L., Li, X. L., Kuncoro, A., Nematzadeh, A., Gribovskaya, E., Donato, D., Lazaridou, A., Mensch, A., Lespiau, J.-B., Tsimpoukelli, M., Grigorev, N., Fritz, D., Sottiaux, T., Pajarskas, M., Pohlen, T., Gong, Z., Toyama, D., de Masson d’Autume, C., Li, Y., Terzi, T., Mikulik, V., Babuschkin, I., Clark, A., de Las Casas, D., Guy, A., Jones, C., Bradbury, J., Johnson, M., Hechtman, B., Weidinger, L., Gabriel, I., Isaac, W., Lockhart, E., Osindero, S., Rimell, L., Dyer, C., Vinyals, O., Ayoub, K., Stanway, J., Bennett, L., Hassabis, D., Kavukcuoglu, K., and Irving, G · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Original
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Later among the works it cites.
Combiner: Full attention Transformer with sparse computation cost
Ren, H., Dai, H., Dai, Z., Yang, M., Leskovec, J., Schuurmans, D., and Dai, B · 2021
Later among the works it cites.
Efficient content-based sparse attention with Routing Transformers
Roy, A., Saffar, M., Vaswani, A., and Grangier, D · 2021
Later among the works it cites.
Image super-resolution via iterative refinement
Original
Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D. J., and Norouzi, M · 2021
Later among the works it cites.
Primer: Searching for efficient Transformers for language modeling
Original
So, D. R., Mańke, W., Liu, H., Dai, Z., Shazeer, N., and Le, Q. V · 2021
Later among the works it cites.
RoFormer: Enhanced Transformer with rotary position embedding
Original
Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y · 2021
Later among the works it cites.
Do long-range language models actually use long-range context?
Sun, S., Krishna, K., Mattarella-Micke, A., and Iyyer, M · 2021
Later among the works it cites.
Mesh transformer Jax, 2021
Wang, B · 2021
Later among the works it cites.
NÜWA: Visual synthesis pre-training for neural visual world creation
Original
Wu, C., Liang, J., Ji, L., Yang, F., Fang, Y., Jiang, D., and Duan, N · 2021
Later among the works it cites.
SoundStream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2021
Later among the works it cites.
Perceiver IO: A general architecture for structured inputs & outputs
Jaegle, A., Borgeaud, S., Alayrac, J.-B., Doersch, C., Ionescu, C., Ding, D., Koppula, S., Zoran, D., Brock, A., Shelhamer, E., Henaff, O., Botvinick, M. M., Zisserman, A., Vinyals, O., and Carreira, J · 2022
Closest in time.
Memorizing transformers
Wu, Y., Rabe, M. N., Hutchins, D., and Szegedy, C · 2022
Closest in time.