Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated remarkable advancements in language understanding and generation.
Y. Xia, F. Tian, L. Wu
2017
Earlier work this paper cites.
M. Shoeybi, M. Patwary, R. Puri
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
B. Zhang and R. Sennrich, “Root mean square layer normalization,”
2019
Earlier work this paper cites.
T. B. Brown, “Language models are few-shot learners,”
2020
Earlier work this paper cites.
K. Hu, T. N. Sainath
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts
2020
Earlier work this paper cites.
B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameter-efficient prompt tuning,”
2021
Earlier work this paper cites.
X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,”
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer, “Rethinking the role of demonstrations: What makes in-context learning work?” in
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans
2022
Earlier work this paper cites.
D. Le, A. Shrivastava, P. Tomasello
2022
Earlier work this paper cites.
A. Chowdhery, S. Narang, J. Devlin
2023
Earlier work this paper cites.
Gemini Team, “Gemini: a family of highly capable multimodal models,”
2023
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal
2023
Cited alongside, same era.
H. Touvron, T. Lavril, G. Izacard
2023
Cited alongside, same era.
2023
Cited alongside, same era.
M. Wang, W. Han, I. Shafran
2023
Cited alongside, same era.
J. Wu, Y. Gaur, Z. Chen
2023
Cited alongside, same era.
C.-H. H. Yang, Y. Gu, Y.-C. Liu, S. Ghosh, I. Bulyko, and A. Stolcke, “Generative speech recognition error correction with large language models and task-activating prompting,” in
Z. Chen, H. Huang, A. Andrusenko
2024
Closest in time.
E. Lakomkin, C. Wu, Y. Fathullah
2024
Closest in time.
X. Yang, W. Kang, Z. Yao
2024
Closest in time.
2024
Closest in time.
M. Reid, N. Savinov, D. Teplyashin
2024
Closest in time.
A. Dubey, A. Jauhri, A. Pandey
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Y. Gong, A. H. Liu, H. Luo
2023
Cited alongside, same era.
A. Radford, J. W. Kim, T. Xu
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Q. Chen, Y. Chu, Z. Gao, Z. Li, K. Hu, X. Zhou, J. Xu, Z. Ma, W. Wang, S. Zheng
2023
Cited alongside, same era.
A. Conneau, M. Ma, S. Khanuja
2023
Cited alongside, same era.
S. Wang, C.-H. Yang, J. Wu, and C. Zhang, “Can whisper perform speech-based in-context learning?” in
2024
Closest in time.
Y. Hu, C. Chen, C.-H. H. Yang, R. Li, D. Zhang, Z. Chen, and E. S. Chng, “Gentranslate: Large language models are generative multilingual speech and machine translators,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
NVIDIA. NVIDIA Riva Megatron NMT any-to-any 1B
2024
Closest in time.
NVIDIA, “Megatron-NMT: Any-to-English, 500M Model,”
2024
Closest in time.
——, “Megatron-NMT: English-to-Any, 500M Model,”
2024
Closest in time.