Fetching the paper…
Reading the bibliography…
LLMs have revolutionized the field of artificial intelligence and have emerged as the de-facto tool for many tasks.
Text and Context: Explorations in the Semantics and Pragmatics of Discourse
V. Dijk · 1977
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Toward better storylines with sentence-level language models
D. Ippolito, D. Grangier, D. Eck, and C. Callison-Burch · 2005
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2006
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2010
Earlier work this paper cites.
Teaching machines to read and comprehend
K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom · 2015
Earlier work this paper cites.
Improved residual vector quantization for high-dimensional approximate nearest neighbor search
S. Liu, H. Lu, and J. Shao · 2015
Earlier work this paper cites.
A corpus and cloze evaluation for deeper understanding of commonsense stories
N. Mostafazadeh, N. Chambers, X. He, D. Parikh, D. Batra, L. Vanderwende, P. Kohli, and J. Allen · 2016
Earlier work this paper cites.
Accurate, large minibatch sg d: training imagenet in 1 hour
P. Goyal · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Effective parallel corpus mining using bilingual sentence embeddings
M. Guo, Q. Shen, Y. Yang, H. Ge, D. Cer, G. H. Abrego, K. Stevens, N. Constant, Y.-H. Sung, B. Strope, et al · 2018
Earlier work this paper cites.
S. Narayan, S. B. Cohen, and M. Lapata · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Earlier work this paper cites.
Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond
M. Artetxe and H. Schwenk · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese Bert-networks
N. Reimers and I. Gurevych · 2019
Earlier work this paper cites.
Neural network acceptability judgments
A. Warstadt, A. Singh, and S. R. Bowman · 2019
Earlier work this paper cites.
Neural text generation with unlikelihood training
S. Welleck, I. Kulikov, S. Roller, E. Dinan, K. Cho, and J. Weston · 2019
Earlier work this paper cites.
Y. Yang, G. H. Abrego, S. Yuan, M. Guo, Q. Shen, D. Cer, Y.-H. Sung, B. Strope, and R. Kurzweil · 2019
Earlier work this paper cites.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov · 2020
Earlier work this paper cites.
Language-agnostic BERT sentence embedding
F. Feng, Y. Yang, D. Cer, N. Arivazhagan, and W. Wang · 2020
Earlier work this paper cites.
spaCy: Industrial-strength Natural Language Processing in Python, 2020
M. Honnibal, I. Montani, S. Van Landeghem, and A. Boyd · 2020
Earlier work this paper cites.
INSET: Sentence infilling with inter-sentential transformer
Y. Huang, Y. Zhang, O. Elachqar, and Y. Cheng · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
N. Kitaev, Ł. Kaiser, and A. Levskaya · 2020
Earlier work this paper cites.
Reformulating unsupervised style transfer as paraphrase generation
K. Krishna, J. Wieting, and M. Iyyer · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Earlier work this paper cites.
W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Y.-A. Chung, Y. Zhang, W. Han, C.-C. Chiu, J. Qin, R. Pang, and Y. Wu · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Earlier work this paper cites.
Multimodal and multilingual embeddings for large-scale speech mining
P.-A. Duquenne, H. Gong, and H. Schwenk · 2021
Earlier work this paper cites.
Using BERT encoding and sentence-level language model for sentence ordering
M. Golestani, S. Z. Razavi, Z. Borhanifard, F. Tahmasebian, and H. Faili · 2021
Earlier work this paper cites.
XL-sum: Large-scale multilingual abstractive summarization for 44 languages
T. Hasan, A. Bhattacharjee, M. S. Islam, K. Mubasshir, Y.-F. Li, Y.-B. Kang, M. S. Rahman, and R. Shahriyar · 2021
Earlier work this paper cites.
Sentence-level planning for especially abstractive summarization
A. Marfurt and J. Henderson · 2021
Earlier work this paper cites.
Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models
J. Ni, G. H. Abrego, N. Constant, J. Ma, K. B. Hall, D. Cer, and Y. Yang · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
A. Q. Nichol and P. Dhariwal · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi · 2021
Cited alongside, same era.
XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y. Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli · 2022
Cited alongside, same era.
mslam: Massively multilingual joint pre-training for speech and text
A. Bapna, C. Cherry, Y. Zhang, Y. Jia, M. Johnson, Y. Cheng, S. Khanuja, J. Riesa, and A. Conneau · 2022
Cited alongside, same era.
Neural codec language models are zero-shot text to speech synthesizers
C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, L. He, S. Zhao, and F. Wei · 2023
Later among the works it cites.
Planner: Generating diversified paragraph via latent language diffusion model
Y. Zhang, J. Gu, Z. Wu, S. Zhai, J. Susskind, and N. Jaitly · 2023
Later among the works it cites.
H. An, Y. Chen, Z. Sun, and X. Li · 2024
Closest in time.
The Claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic · 2024
Closest in time.
Aya 23: Open weight releases to further multilingual progress
V. Aryabumi, J. Dang, D. Talupuru, S. Dash, D. Cairuz, H. Lin, B. Venkitesh, M. Smith, K. Marchisio, S. Ruder, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T-modules: Translation modules for zero-shot cross-modal machine translation
P.-A. Duquenne, H. Gong, B. Sagot, and H. Schwenk · 2022
Cited alongside, same era.
Make-a-scene: Scene-based text-to-image generation with human priors
O. Gafni, A. Polyak, O. Ashual, S. Sheynin, D. Parikh, and Y. Taigman · 2022
Cited alongside, same era.
Problem-solving recognition in scientific text
K. Heffernan and S. Teufel · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, and T. Salimans · 2022
Cited alongside, same era.
Rethinking self-supervision objectives for generalizable coherence modeling
P. Jwalapuram, S. Joty, and X. Lin · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
T. Karras, M. Aittala, T. Aila, and S. Laine · 2022
Cited alongside, same era.
Closest in time.
Self-supervised learning from images with a joint-embedding predictive architecture
M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. Rabbat, Y. LeCun, and N. Ballas · 2024
Closest in time.
Fauno: The italian large language model that will leave you senza parole!
A. Bacciu, G. Trappolini, A. Santilli, E. Rodolà, and F. Silvestri · 2024
Closest in time.
Revisiting feature prediction for learning visual representations from video
A. Bardes, Q. Garrido, J. Ponce, X. Chen, M. Rabbat, Y. LeCun, M. Assran, and N. Ballas · 2024
Closest in time.
ALLaM: Large language models for arabic and english
M. S. Bari, Y. Alnumay, N. A. Alzahrani, N. M. Alotaibi, H. A. Alyahya, S. AlRashed, F. A. Mirza, S. Z. Alsubaie, H. A. Alahmed, G. Alabduljabbar, R. Alkhathran, Y. Almushayqih, R. Alnajim, S. Alsubaihi, M. A. Mansour, M. Alrubaian, A. Alammari, Z. Alawami, A. Al-Thubaity, A. Abdelali, J. Kuriakose, A. Abujabal, N. Al-Twairesh, A. Alowisheq, and H. Khan · 2024
Closest in time.
Cosmopedia, 2024
L. Ben Allal, A. Lozhkov, G. Penedo, T. Wolf, and L. von Werra · 2024
Closest in time.
Chameleon: Mixed-modal early-fusion foundation models
Chameleon team · 2024
Closest in time.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2024
Closest in time.
Learning to plan for language modeling from unlabeled data
N. Cornille, M.-F. Moens, and F. Mai · 2024
Closest in time.
LCFO: Long context and long form output dataset and benchmarking
M. R. Costa-jussà, P. Andrews, M. C. Megliogli, J. Chen, J. Chuang, D. Dale, C. Ropers, A. Mourachko, E. Sánchez, H. Schwenk, T. Tran, A. Turkatenko, and C. Wood · 2024
Closest in time.
M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, P.-E. Mazaré, M. Lomeli, L. Hosseini, and H. Jégou · 2024
Closest in time.
Segment Any Text: A universal approach for robust, efficient and adaptable sentence segmentation
M. Frohmann, I. Sterner, I. Vulić, B. Minixhofer, and M. Schedl · 2024
Closest in time.
Gemini 1.5 unlocking multimodal understanding across millions of tokens of conte
Gemini Team Google · 2024
Closest in time.
Mexma: Token-level objectives improve sentence representations, 2024
J. M. Janeiro, B. Piwowarski, P. Gallinari, and L. Barrault · 2024
Closest in time.
Emma-500: Enhancing massively multilingual adaptation of large language models
S. Ji, Z. Li, I. Paul, J. Paavola, P. Lin, P. Chen, D. O’Brien, H. Luo, H. Schütze, J. Tiedemann, et al · 2024
Closest in time.
Understanding diffusion objectives as the elbo with simple data augmentation
D. Kingma and R. Gao · 2024
Closest in time.
Y. Li, Q. Chen, W. Yan, W. Wang, Q. Zhang, and H. Sundaram · 2024
Closest in time.
Common diffusion noise schedules and sample steps are flawed
S. Lin, B. Liu, J. Li, and X. Yang · 2024
Closest in time.
Diffusion guided language modeling
J. Lovelace, V. Kishore, Y. Chen, and K. Q. Weinberger · 2024
Closest in time.
Fineweb-edu, May 2024
A. Lozhkov, L. Ben Allal, L. von Werra, and T. Wolf · 2024
Closest in time.
Multi-task contrastive learning for 8192-token bilingual text embeddings, 2024
I. Mohr, M. Krimmel, S. Sturua, M. K. Akram, A. Koukounas, M. Günther, G. Mastrapas, V. Ravishankar, J. F. Martínez, F. Wang, Q. Liu, Z. Yu, J. Fu, S. Ognawala, S. Guzman, B. Wang, M. Werk, N. Wang, and H. Xiao · 2024
Closest in time.
OpenAI · 2024
Closest in time.
Tencdm: Understanding the properties of diffusion model in the space of language model encodings
A. Shabalin, V. Meshchaninov, E. Chimbulatov, V. Lapikov, R. Kim, G. Bartosh, D. Molchanov, S. Markov, and D. Vetrov · 2024
Closest in time.
Lola–an open-source massively multilingual large language model
N. Srivastava, D. Kuchelev, T. M. Ngoli, K. Shetty, M. Röder, D. Moussallem, H. Zahera, and A.-C. N. Ngomo · 2024
Closest in time.
jina-embeddings-v3: Multilingual embeddings with task lora, 2024
S. Sturua, I. Mohr, M. K. Akram, M. Günther, B. Wang, M. Krimmel, F. Wang, G. Mastrapas, A. Koukounas, N. Wang, and H. Xiao · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2024
Closest in time.
ChatGLM: A family of large language models from glm-130b to glm-4 all tools
Team GLM, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Zhang, D. Rojas, G. Feng, H. Zhao, H. Lai, H. Yu, H. Wang, J. Sun, J. Zhang, J. Cheng, J. Gui, J. Tang, J. Zhang, J. Sun, J. Li, L. Zhao, L. Wu, L. Zhong, M. Liu, M. Huang, P. Zhang, Q. Zheng, R. Lu, S. Duan, S. Zhang, S. Cao, S. Yang, W. L. Tam, W. Zhao, X. Liu, X. Xia, X. Zhang, X. Gu, X. Lv, X. Liu, X. Liu, X. Yang, X. Song, X. Zhang, Y. An, Y. Xu, Y. Niu, Y. Yang, Y. Li, Y. Bai, Y. Dong, Z. Qi, Z. Wang, Z. Yang, Z. Du, Z. Hou, and Z. Wang · 2024
Closest in time.
The Llama3 team · 2024
Closest in time.
Diffusion model for planning: A systematic literature review
T. Ubukata, J. Li, and K. Tei · 2024
Closest in time.
Aya model: An instruction finetuned open-access multilingual language model
A. Üstün, V. Aryabumi, Z.-X. Yong, W.-Y. Ko, D. D’souza, G. Onilude, N. Bhandari, S. Singh, H.-L. Ooi, A. Kayid, et al · 2024
Closest in time.
Improving text embeddings with large language models
L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, and F. Wei · 2024
Closest in time.
Beyond autoregression: Discrete diffusion for complex reasoning and planning
J. Ye, J. Gao, S. Gong, L. Zheng, X. Jiang, Z. Li, and L. Kong · 2024
Closest in time.
Semformer: Transformer language models with semantic planning
Y. Yin, J. Ding, K. Song, and Y. Zhang · 2024
Closest in time.
X. Zhang, Y. Zhang, D. Long, W. Xie, Z. Dai, J. Tang, H. Lin, B. Yang, P. Xie, F. Huang, M. Zhang, W. Li, and M. Zhang · 2024
Closest in time.