Fetching the paper…
Reading the bibliography…
Self-supervised large language models have demonstrated the ability to perform Machine Translation (MT) via in-context learning, but little is known about where the model performs the task with respect to prompt instructions and demonstration examples.
Voita, E., Sennrich, R., and Titov, I · 1909
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Koehn, P · 2005
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C. J., Mnih, A., and Teh, Y. W · 2016
Earlier work this paper cites.
Learning sparse neural networks through l _ 0 l\_0 regularization
Louizos, C., Welling, M., and Kingma, D. P · 2017
Earlier work this paper cites.
A call for clarity in reporting bleu scores
Post, M · 2018
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
Hewitt, J. and Liang, P · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Earlier work this paper cites.
Losing heads in the lottery: Pruning transformer attention in neural machine translation
Behnke, M. and Heafield, K · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
How do decisions emerge across layers in neural models? interpretation with differentiable masking
De Cao, N., Schlichtkrull, M. S., Aziz, W., and Titov, I · 2020
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Pruning neural machine translation for speed using group lasso
Behnke, M. and Heafield, K · 2021
Earlier work this paper cites.
GPT-Neo: Large scale autoregressive language modeling with Mesh-Tensorflow, March 2021
Black, S., Leo, G., Wang, P., Leahy, C., and Biderman, S · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
The flores-101 evaluation benchmark for low-resource and multilingual machine translation
Goyal, N., Gao, C., Chaudhary, V., Chen, P.-J., Wenzek, G., Ju, D., Krishnan, S., Ranzato, M., Guzmán, F., and Fan, A · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Layer-wise analysis of a self-supervised speech representation model
Pasad, A., Chou, J.-C., and Livescu, K · 2021
Cited alongside, same era.
Fine-tuned transformers show clusters of similar representations across layers
Phang, J., Liu, H., and Bowman, S. R · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2021
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., brian ichter, Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Later among the works it cites.
Xie, S., Qiu, J., Pasad, A., Du, L., Qu, Q., and Mei, H · 2022
Later among the works it cites.
Rethinking the role of scale for in-context learning: An interpretability-based case study at 66 billion scale
Bansal, H., Gopalakrishnan, K., Dingliwal, S., Bodapati, S., Kirchhoff, K., and Roth, D · 2023
Later among the works it cites.
Tart: A plug-and-play transformer module for task-agnostic reasoning, 2023
Bhatia, K., Narayan, A., Sa, C. D., and Ré, C · 2023
Later among the works it cites.
Why can gpt learn in-context? language models implicitly perform gradient descent as meta-optimizers
Dai, D., Sun, Y., Dong, L., Hao, Y., Ma, S., Sui, Z., and Wei, F · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
In-context examples selection for machine translation
Agrawal, S., Zhou, C., Lewis, M., Zettlemoyer, L., and Ghazvininejad, M · 2022
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2022
Cited alongside, same era.
Nearest class-center simplification through intermediate layers
Ben-Shaul, I. and Dekel, S · 2022
Cited alongside, same era.
On the transformation of latent space in fine-tuned nlp models
Durrani, N., Sajjad, H., Dalvi, F., and Alam, F · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P. S., and Valiant, G · 2022
Cited alongside, same era.
The bigscience roots corpus: A 1.6 tb composite multilingual dataset
Laurençon, H., Saulnier, L., Wang, T., Akiki, C., Villanova del Moral, A., Le Scao, T., Von Werra, L., Mou, C., González Ponferrada, E., Nguyen, H., et al · 2022
Cited alongside, same era.
Few-shot learning with multilingual generative language models
Lin, X. V., Mihaylov, T., Artetxe, M., Wang, T., Chen, S., Simig, D., Ott, M., Goyal, N., Bhosale, S., Du, J., Pasunuru, R., Shleifer, S., Koura, P. S., Chaudhary, V., O’Horo, B., Wang, J., Zettlemoyer, L., Kozareva, Z., Diab, M., Stoyanov, V., and Li, X · 2022
Cited alongside, same era.
Later among the works it cites.
A practical survey on faster and lighter transformers
Fournier, Q., Caron, G. M., and Aloise, D · 2023
Later among the works it cites.
The unreasonable effectiveness of few-shot learning for machine translation
Garcia, X., Bansal, Y., Cherry, C., Foster, G., Krikun, M., Feng, F., Johnson, M., and Firat, O · 2023
Later among the works it cites.
How good are gpt models at machine translation? a comprehensive evaluation
Hendy, A., Abdelrehim, M., Sharaf, A., Raunak, V., Gabr, M., Matsushita, H., Kim, Y. J., Afify, M., and Awadalla, H. H · 2023
Later among the works it cites.
The closeness of in-context learning and weight shifting for softmax regression
Li, S., Song, Z., Xia, Y., Yu, T., and Zhou, T · 2023
Later among the works it cites.
Adaptive machine translation with large language models
Moslem, Y., Haque, R., and Way, A · 2023
Later among the works it cites.
On the effect of dropping layers of pre-trained transformer models
Sajjad, H., Dalvi, F., Durrani, N., and Nakov, P · 2023
Later among the works it cites.
Sia, S. and Duh, K · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent, 2023
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
Larger language models do in-context learning differently, 2023
Wei, J., Wei, J., Tay, Y., Tran, D., Webson, A., Lu, Y., Chen, X., Liu, H., Huang, D., Zhou, D., and Ma, T · 2023
Later among the works it cites.
The learnability of in-context learning
Wies, N., Levine, Y., and Shashua, A · 2023
Later among the works it cites.