Fetching the paper…
Reading the bibliography…
Large language models are built on top of a transformer-based architecture to process textual inputs.
Dalal, N., Triggs, B.: Histograms of oriented gradients for human detection. In: 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05) (2005)
2005
Earlier work this paper cites.
Hyv"arinen, A., Dayan, P.: Estimation of non-normalized statistical models by score matching. J. Mach. Learn. Res. (2005)
2005
Earlier work this paper cites.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: Proc. IEEE Conf. Comp. Vis. Patt. Recogn. (2009)
2009
Earlier work this paper cites.
Kingma, D.P., Welling, M.: Auto-encoding variational bayes. arXiv: Comp. Res. Repository (2013)
2013
Earlier work this paper cites.
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial nets. In: Proc. Advances in Neural Inf. Process. Syst. (2014)
2014
Earlier work this paper cites.
Ronneberger, O., Fischer, P., Brox, T.: U-net: Convolutional networks for biomedical image segmentation. In: Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18 (2015)
2015
Earlier work this paper cites.
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., Ganguli, S.: Deep unsupervised learning using nonequilibrium thermodynamics. In: Proc. Int. Conf. Mach. Learn. (2015)
2015
Earlier work this paper cites.
Ba, J.L., Kiros, J.R., Hinton, G.E.: Layer normalization. arXiv: Comp. Res. Repository (2016)
2016
Earlier work this paper cites.
Huang, G., Sun, Y., Liu, Z., Sedra, D., Weinberger, K.Q.: Deep networks with stochastic depth. In: Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14 (2016)
2016
Earlier work this paper cites.
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., Chen, X.: Improved techniques for training gans. Proc. Advances in Neural Inf. Process. Syst. (2016)
2016
Earlier work this paper cites.
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., Wojna, Z.: Rethinking the inception architecture for computer vision. In: Proc. IEEE Conf. Comp. Vis. Patt. Recogn. (2016)
2016
Earlier work this paper cites.
He, K., Gkioxari, G., Dollár, P., Girshick, R.: Mask r-cnn. In: Proc. IEEE Int. Conf. Comp. Vis. (2017)
2017
Earlier work this paper cites.
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S.: Gans trained by a two time-scale update rule converge to a local nash equilibrium. Proc. Advances in Neural Inf. Process. Syst. (2017)
2017
Earlier work this paper cites.
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. arXiv: Comp. Res. Repository (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. Proc. Advances in Neural Inf. Process. Syst. (2017)
2017
Earlier work this paper cites.
You, Y., Gitman, I., Ginsburg, B.: Large batch training of convolutional networks. arXiv: Comp. Res. Repository (2017)
2017
Earlier work this paper cites.
Zhang, H., Cisse, M., Dauphin, Y.N., Lopez-Paz, D.: mixup: Beyond empirical risk minimization. arXiv: Comp. Res. Repository (2017)
2017
Earlier work this paper cites.
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., Torralba, A.: Scene parsing through ade20k dataset. In: Proc. IEEE Conf. Comp. Vis. Patt. Recogn. (2017)
2017
Earlier work this paper cites.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv: Comp. Res. Repository (2018)
2018
Earlier work this paper cites.
Shaw, P., Uszkoreit, J., Vaswani, A.: Self-attention with relative position representations. arXiv: Comp. Res. Repository (2018)
2018
Earlier work this paper cites.
Xiao, T., Liu, Y., Zhou, B., Jiang, Y., Sun, J.: Unified perceptual parsing for scene understanding. In: Proc. Eur. Conf. Comp. Vis. (2018)
2018
Earlier work this paper cites.
Kynkäänniemi, T., Karras, T., Laine, S., Lehtinen, J., Aila, T.: Improved precision and recall metric for assessing generative models. Proc. Advances in Neural Inf. Process. Syst. (2019)
2019
Earlier work this paper cites.
Yun, S., Han, D., Oh, S.J., Chun, S., Choe, J., Yoo, Y.: Cutmix: Regularization strategy to train strong classifiers with localizable features. In: Proc. IEEE Int. Conf. Comp. Vis. (2019)
2019
Earlier work this paper cites.
Zhang, B., Sennrich, R.: Root mean square layer normalization. Proc. Advances in Neural Inf. Process. Syst. (2019)
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language models are few-shot learners. Proc. Advances in Neural Inf. Process. Syst. (2020)
2020
Earlier work this paper cites.
Contributors, M.: Mmsegmentation: Openmmlab semantic segmentation toolbox and benchmark (2020)
2020
Earlier work this paper cites.
Cubuk, E.D., Zoph, B., Shlens, J., Le, Q.V.: Randaugment: Practical automated data augmentation with a reduced search space. In: Proc. IEEE Conf. Comp. Vis. Patt. Recogn. (2020)
2020
Earlier work this paper cites.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. arXiv: Comp. Res. Repository (2020)
2020
Cited alongside, same era.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. Proc. Advances in Neural Inf. Process. Syst. (2020)
2020
Cited alongside, same era.
Shazeer, N.: Glu variants improve transformer. arXiv: Comp. Res. Repository (2020)
2020
Cited alongside, same era.
Song, Y., Sohl-Dickstein, J., Kingma, D.P., Kumar, A., Ermon, S., Poole, B.: Score-based generative modeling through stochastic differential equations. arXiv: Comp. Res. Repository (2020)
2020
Cited alongside, same era.
Bao, H., Dong, L., Piao, S., Wei, F.: Beit: Bert pre-training of image transformers. arXiv: Comp. Res. Repository (2021)
Chu, X., Qiao, L., Lin, X., Xu, S., Yang, Y., Hu, Y., Wei, F., Zhang, X., Zhang, B., Wei, X., et al.: Mobilevlm: A fast, reproducible and strong vision language assistant for mobile devices. arXiv: Comp. Res. Repository (2023)
2023
Later among the works it cites.
Chu, X., Tian, Z., Zhang, B., Wang, X., Shen, C.: Conditional positional encodings for vision transformers. In: Proc. Int. Conf. Learn. Representations (2023)
2023
Later among the works it cites.
Contributors, M.: Openmmlab’s pre-training toolbox and benchmark (2023)
2023
Later among the works it cites.
Dao, T.: Flashattention-2: Faster attention with better parallelism and work partitioning. arXiv: Comp. Res. Repository (2023)
2023
Later among the works it cites.
Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., Jia, J.: Lisa: Reasoning segmentation via large language model. arXiv: Comp. Res. Repository (2023)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
Chu, X., Tian, Z., Wang, Y., Zhang, B., Ren, H., Wei, X., Xia, H., Shen, C.: Twins: Revisiting the design of spatial attention in vision transformers. In: Proc. Advances in Neural Inf. Process. Syst. (2021)
2021
Cited alongside, same era.
Dhariwal, P., Nichol, A.: Diffusion models beat gans on image synthesis. Proc. Advances in Neural Inf. Process. Syst. (2021)
2021
Cited alongside, same era.
Ho, J., Salimans, T.: Classifier-free diffusion guidance. In: NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications (2021)
2021
Cited alongside, same era.
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: Swin transformer: Hierarchical vision transformer using shifted windows. In: Proc. IEEE Int. Conf. Comp. Vis. (2021)
2021
Cited alongside, same era.
Nash, C., Menick, J., Dieleman, S., Battaglia, P.W.: Generating images with sparse representations. arXiv: Comp. Res. Repository (2021)
2021
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: Proc. Int. Conf. Mach. Learn. (2021)
2021
Cited alongside, same era.
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., Jegou, H.: Training data-efficient image transformers and distillation through attention. In: Proc. Int. Conf. Mach. Learn. (2021)
2021
Cited alongside, same era.
2023
Later among the works it cites.
Liu, H., Li, C., Li, Y., Lee, Y.J.: Improved baselines with visual instruction tuning. arXiv: Comp. Res. Repository (2023)
2023
Later among the works it cites.
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. Pattern Recogn. (2023)
2023
Later among the works it cites.
Liu, Y., Zhang, S., Chen, J., Yu, Z., Chen, K., Lin, D.: Improving pixel-based mim by reducing wasted modeling capability. In: Proc. IEEE Int. Conf. Comp. Vis. (2023)
2023
Later among the works it cites.
OpenAI: Gpt-4 technical report (2023)
2023
Later among the works it cites.
Peebles, W., Xie, S.: Scalable diffusion models with transformers. In: Proc. IEEE Int. Conf. Comp. Vis. (2023)
2023
Later among the works it cites.
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X.E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al.: Code llama: Open foundation models for code. arXiv: Comp. Res. Repository (2023)
2023
Later among the works it cites.
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., Liu, Y.: Roformer: Enhanced transformer with rotary position embedding. Int. J. Comput. Vision (2023)
2023
Later among the works it cites.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F.: Llama: Open and efficient foundation language models. arXiv: Comp. Res. Repository (2023)
2023
Later among the works it cites.
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al.: Llama 2: Open foundation and fine-tuned chat models. arXiv: Comp. Res. Repository (2023)
2023
Later among the works it cites.
Vishniakov, K., Shen, Z., Liu, Z.: Convnet vs transformer, supervised vs clip: Beyond imagenet accuracy. arXiv: Comp. Res. Repository (2023)
2023
Later among the works it cites.
Wei, F., Zhang, X., Zhang, A., Zhang, B., Chu, X.: Lenna: Language enhanced reasoning detection assistant. arXiv: Comp. Res. Repository (2023)
2023
Later among the works it cites.
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., Han, S.: Smoothquant: Accurate and efficient post-training quantization for large language models. In: Proc. Int. Conf. Mach. Learn. (2023)
2023
Later among the works it cites.
Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K.A., Oguz, B., et al.: Effective long-context scaling of foundation models. arXiv: Comp. Res. Repository (2023)
2023
Later among the works it cites.
Yang, A., Xiao, B., Wang, B., Zhang, B., Bian, C., Yin, C., Lv, C., Pan, D., Wang, D., Yan, D., et al.: Baichuan 2: Open large-scale language models. arXiv: Comp. Res. Repository (2023)
2023
Later among the works it cites.
Brooks, T., Peebles, B., Homes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C.W.Y., Wang, R., Ramesh, A.: Video generation models as world simulators (2024)
2024
Closest in time.
Chen, X., Liu, Z., Xie, S., He, K.: Deconstructing denoising diffusion models for self-supervised learning. arXiv: Comp. Res. Repository (2024)
2024
Closest in time.
Chu, X., Qiao, L., Zhang, X., Xu, S., Wei, F., Yang, Y., Sun, X., Hu, Y., Lin, X., Zhang, B., et al.: Mobilevlm v2: Faster and stronger baseline for vision language model. arXiv: Comp. Res. Repository (2024)
2024
Closest in time.
https://stability.ai/: Stable code 3b: Coding on the edge (2024)
2024
Closest in time.
Li, L., Li, Q., Zhang, B., Chu, X.: Norm tweaking: High-performance low-bit quantization of large language models. In: Proc. AAAI Conf. Artificial Intell. (2024)
2024
Closest in time.
Lu, Z., Wang, Z., Huang, D., Wu, C., Liu, X., Ouyang, W., Bai, L.: Fit: Flexible vision transformer for diffusion model. arXiv: Comp. Res. Repository (2024)
2024
Closest in time.
Ma, N., Goldstein, M., Albergo, M.S., Boffi, N.M., Vanden-Eijnden, E., Xie, S.: Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers. arXiv: Comp. Res. Repository (2024)
2024
Closest in time.
Zhu, D., Chen, J., Shen, X., Li, X., Elhoseiny, M.: MiniGPT-4: Enhancing vision-language understanding with advanced large language models. In: Proc. Int. Conf. Learn. Representations (2024)
2024
Closest in time.