Fetching the paper…
Reading the bibliography…
Recent studies have revealed various manifestations of position bias in transformer architectures, from the "lost-in-the-middle" phenomenon to attention sinks, yet a comprehensive theoretical understanding of how attention masks and positional encodings shape these biases remains elusive.
Two storage mechanisms in free recall
Glanzer, M. and Cunitz, A. R · 1966
Earlier work this paper cites.
An Introduction to Functional Grammar
Halliday, M. A · 2004
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Structured attention networks
Kim, Y., Denton, C., Hoang, L., and Rush, A. M · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N. M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
How contextual are contextualized word representations? comparing the geometry of bert, elmo, and gpt-2 embeddings
Ethayarajh, K · 2019
Earlier work this paper cites.
Representation degeneration problem in training natural language generation models
Gao, J., He, D., Tan, X., Qin, T., Wang, L., and Liu, T.-Y · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
Quantifying attention flow in transformers
Abnar, S. and Zuidema, W · 2020
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Attention is not only a weight: Analyzing transformers with vector norms
Kobayashi, G., Kuribayashi, T., Yokoi, S., and Inui, K · 2020
Earlier work this paper cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., rahman Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Isotropy in the contextual embedding space: Clusters and manifolds
Cai, X., Huang, J., Bian, Y.-L., and Church, K. W · 2021
Earlier work this paper cites.
Multilingual language models predict human reading behavior
Hollenstein, N., Pirovano, F., Zhang, C., Jäger, L. A., and Beinborn, L · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zhao, T. Z., Wallace, E., Feng, S., Klein, D., and Singh, S · 2021
Cited alongside, same era.
Kerple: Kernelized relative positional embedding for length extrapolation
Chi, T.-C., Fan, T.-H., Ramadge, P. J., and Rudnicky, A. I · 2022
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P · 2022
Cited alongside, same era.
The parallelism tradeoff: Limitations of log-precision transformers
Merrill, W. and Sabharwal, A · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
Anisotropy is inherent to self-attention in transformers
Godey, N., de la Clergerie, E. V., and Sagot, B · 2024
Later among the works it cites.
Active-dormant attention heads: Mechanistically demystifying extreme-token phenomena in llms
Guo, T., Pai, D., Bai, Y., Jiao, J., Jordan, M. I., and Mei, S · 2024
Later among the works it cites.
Serial position effects of large language models
Guo, X. and Vosoughi, S · 2024
Later among the works it cites.
Large language models are zero-shot rankers for recommender systems
Hou, Y., Zhang, J., Lin, Z., Lu, H., Xie, R., McAuley, J., and Zhao, W. X · 2024
Later among the works it cites.
Lost in the middle: How language models use long contexts
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2022
Cited alongside, same era.
Is anisotropy truly harmful? a case study on text clustering
Ait-Saada, M. and Nadif, M · 2023
Cited alongside, same era.
Fast attention requires bounded entries
Alman, J. and Song, Z · 2023
Cited alongside, same era.
Mistral 7b
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Cited alongside, same era.
The impact of positional encoding on length generalization in transformers
Kazemnejad, A., Padhi, I., Ramamurthy, K. N., Das, P., and Reddy, S · 2023
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2023
Cited alongside, same era.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task
Reddy, G · 2024
Later among the works it cites.
One-layer transformers fail to solve the induction heads task
Sanford, C., Hsu, D., and Telgarsky, M · 2024
Later among the works it cites.
Eliminating position bias of language models: A mechanistic approach
Wang, Z., Zhang, H., Li, X., Huang, K.-H., Han, C., Ji, S., Kakade, S. M., Peng, H., and Ji, H · 2024
Later among the works it cites.
On the role of attention masks and layernorm in transformers
Wu, X., Ajorlou, A., Wang, Y., Jegelka, S., and Jadbabaie, A · 2024
Later among the works it cites.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Later among the works it cites.
Mitigate position bias in large language models via scaling a single dimension
Yu, Y., Jiang, H., Luo, X., Wu, Q., Lin, C.-Y., Li, D., Yang, Y., Huang, Y., and Qiu, L · 2024
Later among the works it cites.
Zhang, Z. A., Chen, R., Liu, S., Yao, Z., Ruwase, O., Chen, B., Wu, X., and Wang, Z · 2024
Later among the works it cites.
Rethinking invariance in in-context learning
Fang, L., Yifei Wang, K. G., Fang, L., and Wang, Y · 2025
Closest in time.
When attention sink emerges in language models: An empirical view
Gu, X., Pang, T., Du, C., Liu, Q., Zhang, F., Du, C., Wang, Y., and Lin, M · 2025
Closest in time.