Fetching the paper…
Reading the bibliography…
Parameter generation has long struggled to match the scale of today large vision and language models, curbing its broader utility.
Chaos in random neural networks
Sompolinsky, H., Crisanti, A., and Sommers, H.-J · 1988
Earlier work this paper cites.
Stochastic gradient learning in neural networks
Bottou, L. et al · 1991
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen netzen
Hochreiter, S · 1991
Earlier work this paper cites.
Long short-term memory
Hochreiter, S · 1997
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., Schmidhuber Jürgen, Hochreiter, S., and Schmidhuber · 1997
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Practical variational inference for neural networks
Graves, A · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Bayesian learning for neural networks , volume 118
Neal, R. M · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Hypernetworks
Ha, D., Dai, A. M., and Le, Q. V · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., and Torralba, A · 2017
Earlier work this paper cites.
SMASH: One-shot model architecture search through hypernetworks
Brock, A., Lim, T., Ritchie, J., and Weston, N · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Unified perceptual parsing for scene understanding
Xiao, T., Liu, Y., Zhou, B., Jiang, Y., and Sun, J · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Socialiqa: Commonsense reasoning about social interactions
Sap, M., Rashkin, H., Chen, D., LeBras, R., and Choi, Y · 2019
Cited alongside, same era.
GitHub repository: Pytorch image models
Wightman, R · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Cited alongside, same era.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2020
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2023
Later among the works it cites.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Grella, M., et al · 2023
Later among the works it cites.
Song, Y., Dhariwal, P., Chen, M., and Sutskever, I · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
Neural mechanics: Symmetry and broken conservation laws in deep learning dynamics
Kunin, D., Sagastuy-Brena, J., Ganguli, S., Yamins, D. L., and Tanaka, H · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
Nichol, A. Q. and Dhariwal, P · 2021
Cited alongside, same era.
Later among the works it cites.
Diffusion probabilistic model made slim
Yang, X., Zhou, D., Feng, J., and Wang, X · 2023
Later among the works it cites.
xlstm: Extended long short-term memory
Beck, M., Pöppel, K., Spanring, M., Auer, A., Prudnikova, O., Kopp, M., Klambauer, G., Brandstetter, J., and Hochreiter, S · 2024
Later among the works it cites.
Dao, T. and Gu, A · 2024
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2024
Later among the works it cites.
Conditional lora parameter generation
Jin, X., Wang, K., Tang, D., Zhao, W., Zhou, Y., Tang, J., and You, Y · 2024
Later among the works it cites.
Unleash graph neural networks from heavy tuning
Lin, L., Shi, D., Han, A., Wang, Z., and Gao, J · 2024
Later among the works it cites.
Dora: Weight-decomposed low-rank adaptation
Liu, S.-Y., Wang, C.-Y., Yin, H., Molchanov, P., Wang, Y.-C. F., Cheng, K.-T., and Chen, M.-H · 2024
Later among the works it cites.
Deepcache: Accelerating diffusion models for free
Ma, X., Fang, G., and Wang, X · 2024
Later among the works it cites.
Introducing meta llama 3: The most capable openly available llm to date
Meta, A · 2024
Later among the works it cites.
T-stitch: Accelerating sampling in pre-trained diffusion models with trajectory stitching
Pan, Z., Zhuang, B., Huang, D.-A., Nie, W., Yu, Z., Xiao, C., Cai, J., and Anandkumar, A · 2024
Later among the works it cites.
Towards scalable and versatile weight space learning
Schürholt, K., Mahoney, M. W., and Borth, D · 2024
Later among the works it cites.
Temporal dynamic quantization for diffusion models
So, J., Lee, J., Ahn, D., Kim, H., and Park, E · 2024
Later among the works it cites.
Diffusion-based neural network weights generation
Soro, B., Andreis, B., Lee, H., Chong, S., Hutter, F., and Hwang, S. J · 2024
Later among the works it cites.
Wang, K., Xu, Z., Zhou, Y., Zang, Z., Darrell, T., Liu, Z., and You, Y · 2024
Later among the works it cites.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Tang, J., Wang, J., Yang, J., Tu, J., Zhang, J., Ma, J., Xu, J., Zhou, J., Bai, J., He, J., Lin, J., Dang, K., Lu, K., Chen, K., Yang, K., Li, M., Xue, M., Ni, N., Zhang, P., Wang, P., Peng, R., Men, R., Gao, R., Lin, R., Wang, S., Bai, S., Tan, S., Zhu, T., Li, T., Liu, T., Ge, W., Deng, X., Zhou, X., Ren, X., Zhang, X., Wei, X., Ren, X., Fan, Y., Yao, Y., Zhang, Y., Wan, Y., Chu, Y., Liu, Y., Cui, Z., Zhang, Z., and Fan, Z · 2024
Later among the works it cites.
Dynamic tuning towards parameter and inference efficiency for vit adaptation
Zhao, W., Tang, J., Han, Y., Song, Y., Wang, K., Huang, G., Wang, F., and You, Y · 2024
Later among the works it cites.