Fetching the paper…
Reading the bibliography…
Vision Transformers (ViTs) with self-attention modules have recently achieved great empirical success in many vision tasks.
Attention is All you Need
Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A., Kaiser L., Polosukhin I. · 2017
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford A., Kim J., Hallacy C., Ramesh A., Goh G., Agarwal S., Sastry G., Askell A., Mishkin P., Clark J., Krueger G., Sutskever I. (2021) · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy A., Beyer L., Kolesnikov A., Weissenborn D., Zhai X., Unterthiner T., Dehghani M., Minderer M., Heigold G., Gelly S., Uszkoreit J., Houlsby N · 2021
Cited alongside, same era.
GPT-4 Technical Report
OpenAI (2023) · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…