Fetching the paper…
Reading the bibliography…
We explore different ways to utilize position-based cross-attention in seq2seq networks to enable length generalization in algorithmic tasks.
The eccentricity effect: Target eccentricity affects performance on conjunction searches
Carrasco, M., Evert, D. L., Chang, I., and Katz, S. M · 1995
Earlier work this paper cites.
Natural Language Processing with Python
Bird, S., Klein, E., and Loper, E · 2009
Earlier work this paper cites.
Self-delimiting neural networks
Schmidhuber, J · 2012
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gülçehre, Ç., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Graves, A · 2016
Earlier work this paper cites.
Kaiser, L. and Sutskever, I · 2016
Earlier work this paper cites.
Position Information in Transformers: An Overview
Dufter, P., Schmitt, M., and Schütze, H · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Lake, B. and Baroni, M · 2018
Earlier work this paper cites.
Memorize or generalize? searching for a compositional RNN in a haystack
Liska, A., Kruszewski, G., and Baroni, M · 2018
Earlier work this paper cites.
Modeling localness for self-attention networks
Yang, B., Tu, Z., Wong, D. F., Meng, F., Chao, L. S., and Zhang, T · 2018
Earlier work this paper cites.
Deep equilibrium models
Bai, S., Kolter, J. Z., and Koltun, V · 2019
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q., and Salakhutdinov, R · 2019
Cited alongside, same era.
Universal transformers
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, L · 2019
Cited alongside, same era.
Pay less attention with lightweight and dynamic convolutions
Wu, F., Fan, A., Baevski, A., Dauphin, Y., and Auli, M · 2019
Cited alongside, same era.
Location Attention for Extrapolation to Longer Sequences
Dubois, Y., Dagan, G., Hupkes, D., and Bruni, E · 2020
Cited alongside, same era.
Measuring compositional generalization: A comprehensive method on realistic data
Keysers, D., Schärli, N., Scales, N., Buisman, H., Furrer, D., Kashubin, S., Momchev, N., Sinopalnikov, D., Stafiniak, L., Tihon, T., Tsarkov, D., Wang, X., van Zee, M., and Bousquet, O · 2020
Stable, fast and accurate: Kernelized attention with relative positional encoding
Luo, S., Li, S., Cai, T., He, D., Peng, D., Zheng, S., Ke, G., Wang, L., and Liu, T.-Y · 2021
Later among the works it cites.
Show your work: Scratchpads for intermediate computation with language models
Nye, M. I., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A · 2021
Later among the works it cites.
Explore better relative position embeddings from encoding perspective for transformer models
Qu, A., Niu, J., and Mo, S · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y · 2021
Later among the works it cites.
The case for translation-invariant self-attention in transformer-based language models
Wennberg, U. and Henter, G. E · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2020
Cited alongside, same era.
The eos decision and length extrapolation
Newman, B., Hewitt, J., Liang, P., and Manning, C. D · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Encoding word order in complex embeddings
Wang, B., Zhao, D., Lioma, C., Li, Q., Zhang, P., and Simonsen, J. G · 2020
Cited alongside, same era.
Banino, A., Balaguer, J., and Blundell, C · 2021
Cited alongside, same era.
Convolutions and self-attention: Re-interpreting relative positions in pre-trained language models
Chang, T., Xu, Y., Xu, W., and Tu, Z · 2021
Cited alongside, same era.
Later among the works it cites.
DA-transformer: Distance-aware transformer
Wu, C., Wu, F., and Huang, Y · 2021
Later among the works it cites.
Exploring length generalization in large language models
Anil, C., Wu, Y., Andreassen, A. J., Lewkowycz, A., Misra, V., Ramasesh, V. V., Slone, A., Gur-Ari, G., Dyer, E., and Neyshabur, B · 2022
Later among the works it cites.
The neural data router: Adaptive control flow in transformers improves systematic generalization
Csordás, R., Irie, K., and Schmidhuber, J · 2022
Later among the works it cites.
LISA: Learning interpretable skill abstractions from language
Garg, D., Vaidyanath, S., Kim, K., Song, J., and Ermon, S · 2022
Later among the works it cites.
Kim, N., Linzen, T., and Smolensky, P · 2022
Later among the works it cites.
Your transformer may not be as powerful as you expect
Luo, S., Li, S., Zheng, S., Liu, T.-Y., Wang, L., and He, D · 2022
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N., and Lewis, M · 2022
Later among the works it cites.
A length-extrapolatable transformer
Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., and Wei, F · 2022
Later among the works it cites.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., and Vaswani, A · 2074
Closest in time.