Fetching the paper…
Reading the bibliography…
In this work, we leverage the intrinsic segmentation of language sequences and design a new positional encoding method called Bilevel Positional Encoding (BiPE).
Randomized positional encodings boost length generalization of transformers
Ruoss, A., Delétang, G., Genewein, T., Grau-Moya, J., Csordás, R., Bennani, M., Legg, S., and Veness, J · 1903
Earlier work this paper cites.
Automata, languages, and machines
Eilenberg, S · 1974
Earlier work this paper cites.
The hierarchical hidden markov model: Analysis and applications
Fine, S., Singer, Y., and Tishby, N · 1998
Earlier work this paper cites.
Communicating hierarchical state machines
Alur, R., Kannan, S., and Yannakakis, M · 1999
Earlier work this paper cites.
Latent dirichlet allocation
Blei, D. M., Ng, A. Y., and Jordan, M. I · 2003
Earlier work this paper cites.
Hierarchical topic models and the nested chinese restaurant process
Griffiths, T., Jordan, M., Tenenbaum, J., and Blei, D · 2003
Earlier work this paper cites.
Halliday’s introduction to functional grammar
Halliday, M. A. K. and Matthiessen, C. M · 2013
Earlier work this paper cites.
Hierarchical multiscale recurrent neural networks
Chung, J., Ahn, S., and Bengio, Y · 2016
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
The NarrativeQA reading comprehension challenge
Kočiský, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., and Vaswani, A · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Fine-tune bert for extractive summarization
Liu, Y · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Le Bras, R., Bhagavatula, C., and Choi, Y · 2020
Cited alongside, same era.
Segatron: Segment-aware transformer for language modeling and understanding
Bai, H., Shi, P., Lin, J., Xie, Y., Tan, L., Xiong, K., Gao, W., and Li, M · 2021
Cited alongside, same era.
A dataset of information-seeking questions and answers anchored in research papers
Dasigi, P., Lo, K., Beltagy, I., Cohan, A., Smith, N. A., and Gardner, M · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Efficient attentions for long document summarization
Huang, L., Cao, S., Parulian, N., Ji, H., and Wang, L · 2021
Cited alongside, same era.
ContractNLI: A dataset for document-level natural language inference for contracts
CoLT5: Faster long-range transformers with conditional computation
Ainslie, J., Lei, T., de Jong, M., Ontanon, S., Brahma, S., Zemlyanskiy, Y., Uthus, D., Guo, M., Lee-Thorp, J., Tay, Y., Sung, Y.-H., and Sanghai, S · 2023
Later among the works it cites.
Dissecting transformer length extrapolation via the lens of receptive field analysis
Chi, T.-C., Fan, T.-H., Rudnicky, A., and Ramadge, P · 2023
Later among the works it cites.
Monotonic location attention for length generalization
Chowdhury, J. R. and Caragea, C · 2023
Later among the works it cites.
Towards revealing the mystery behind chain of thought: A theoretical perspective
Feng, G., Zhang, B., Gu, Y., Ye, H., He, D., and Wang, L · 2023
Later among the works it cites.
A framework for few-shot language model evaluation, 12 2023
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Koreeda, Y. and Manning, C · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding, 2021
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2021
Cited alongside, same era.
QMSum: A new benchmark for query-based multi-domain meeting summarization
Zhong, M., Yin, D., Yu, T., Zaidi, A., Mutuma, M., Jha, R., Awadallah, A. H., Celikyilmaz, A., Liu, Y., Qiu, X., and Radev, D · 2021
Cited alongside, same era.
Exploring length generalization in large language models
Anil, C., Wu, Y., Andreassen, A., Lewkowycz, A., Misra, V., Ramasesh, V., Slone, A., Gur-Ari, G., Dyer, E., and Neyshabur, B · 2022
Cited alongside, same era.
SummScreen: A dataset for abstractive screenplay summarization
Chen, M., Chu, Z., Wiseman, S., and Gimpel, K · 2022
Cited alongside, same era.
Kerple: Kernelized relative positional embedding for length extrapolation
Chi, T.-C., Fan, T.-H., Ramadge, P. J., and Rudnicky, A · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways, 2022
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
Cited alongside, same era.
Later among the works it cites.
Lm-infinite: Simple on-the-fly length generalization for large language models, 2023
Han, C., Wang, Q., Xiong, W., Chen, Y., Ji, H., and Wang, S · 2023
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
Kazemnejad, A., Padhi, I., Ramamurthy, K. N., Das, P., and Reddy, S · 2023
Later among the works it cites.
Functional interpolation for relative positions improves long context transformers
Li, S., You, C., Guruganesh, G., Ainslie, J., Ontanon, S., Zaheer, M., Sanghai, S., Yang, Y., Kumar, S., and Bhojanapalli, S · 2023
Later among the works it cites.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Later among the works it cites.
Scaling laws of rope-based extrapolation
Liu, X., Yan, H., Zhang, S., An, C., Qiu, X., and Lin, D · 2023
Later among the works it cites.
Yarn: Efficient context window extension of large language models, 2023
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2023
Later among the works it cites.
Parallel context windows for large language models
Ratner, N., Levine, Y., Belinkov, Y., Ram, O., Magar, I., Abend, O., Karpas, E., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al · 2023
Later among the works it cites.
A length-extrapolatable transformer
Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., and Wei, F · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks, 2023
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Later among the works it cites.
Pose: Efficient context window extension of llms via positional skip-wise training
Zhu, D., Yang, N., Wang, L., Song, Y., Wu, W., Wei, F., and Li, S · 2023
Later among the works it cites.
Llm maybe longlm: Self-extend llm context window without tuning
Jin, H., Han, X., Yang, J., Jiang, Z., Liu, Z., Chang, C.-Y., Chen, H., and Hu, X · 2024
Closest in time.