Fetching the paper…
Reading the bibliography…
Transformers are widely used to extract semantic meanings from input tokens, yet they usually operate as black-box models.
Uncertainty principles and signal recovery
Donoho, D. L. and Stark, P. B · 1989
Earlier work this paper cites.
Optimally sparse representation in general (nonorthogonal) dictionaries via ℓ 1 \ell^{1} minimization
Donoho, D. L. and Elad, M · 2003
Earlier work this paper cites.
Decoding by linear programming
Candes, E. J. and Tao, T · 2005
Earlier work this paper cites.
Robust uncertainty principles: Exact signal reconstruction from highly incomplete frequency information
Candès, E. J., Romberg, J., and Tao, T · 2006
Earlier work this paper cites.
Compressed sensing
Donoho, D. L · 2006
Earlier work this paper cites.
On model selection consistency of lasso
Zhao, P. and Yu, B · 2006
Earlier work this paper cites.
Sampling from large matrices: An approach through geometric functional analysis
Rudelson, M. and Vershynin, R · 2007
Earlier work this paper cites.
Introduction to Fourier analysis and wavelets , volume 102
Pinsky, M. A · 2008
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Discrete Fourier analysis and wavelets: applications to signal and image processing
Broughton, S. A. and Bryan, K · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Beyond word importance: Contextual decomposition to extract interactions from lstms
Murdoch, W. J., Liu, P. J., and Yu, B · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., and Vaswani, A · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science , volume 47
Vershynin, R · 2018
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Ethayarajh, K · 2019
Earlier work this paper cites.
Representation degeneration problem in training natural language generation models
Gao, J., He, D., Tan, X., Qin, T., Wang, L., and Liu, T.-Y · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
Hewitt, J. and Manning, C. D · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Visualizing and measuring the geometry of bert
Reif, E., Yuan, A., Wattenberg, M., Viegas, F. B., Coenen, A., Pearce, A., and Kim, B · 2019
Cited alongside, same era.
Transformer dissection: a unified understanding of transformer’s attention via the lens of kernel
Tsai, Y.-H. H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdinov, R · 2019
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2021
Later among the works it cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2022
Later among the works it cites.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vig, J. and Belinkov, Y · 2019
Cited alongside, same era.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Isotropy in the contextual embedding space: Clusters and manifolds
Cai, X., Huang, J., Bian, Y., and Church, K · 2020
Cited alongside, same era.
Finding universal grammatical relations in multilingual bert
Chi, E. A., Hewitt, J., and Manning, C. D · 2020
Cited alongside, same era.
Rethinking positional encoding in language pre-training
Ke, G., He, D., and Liu, T.-Y · 2020
Cited alongside, same era.
Topic modeling with contextualized word representation clusters
Thompson, L. and Mimno, D · 2020
Cited alongside, same era.
Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Later among the works it cites.
Transformers learn through gradual rank increase
Boix-Adsera, E., Littwin, E., Abbe, E., Bengio, S., and Susskind, J · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Closest in time.
Screenot: Exact mse-optimal singular value thresholding in correlated noise
Donoho, D., Gavish, M., and Romanov, E · 2023
Closest in time.
The impact of positional encoding on length generalization in transformers
Kazemnejad, A., Padhi, I., Ramamurthy, K. N., Das, P., and Reddy, S · 2023
Closest in time.
Teaching arithmetic to small transformers
Lee, N., Sreenivasan, K., Lee, J. D., Lee, K., and Papailiopoulos, D · 2023
Closest in time.
Task arithmetic in the tangent space: Improved editing of pre-trained models
Ortiz-Jimenez, G., Favero, A., and Frossard, P · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Mimetic initialization of self-attention layers
Trockman, A. and Kolter, J. Z · 2023
Closest in time.
Absolute position embedding learns sinusoid-like waves for attention based on relative position
Yamamoto, Y. and Matsuzaki, T · 2023
Closest in time.
Attentionviz: A global view of transformer attention
Yeh, C., Chen, Y., Wu, A., Chen, C., Viégas, F., and Wattenberg, M · 2023
Closest in time.