Fetching the paper…
Reading the bibliography…
Structured (dictionary-like) data presents challenges for left-to-right language models, as they can struggle with structured entities for a wide variety of reasons such as formatting and sensitivity to the order in which attributes are presented.
Synthesis of the elements in stars
Burbidge, M. E., Burbidge, G. R., Fowler, W. A., and Hoyle, F · 1957
Earlier work this paper cites.
Neutron radii in nuclei and the neutron equation of state
Brown, B. A · 2000
Earlier work this paper cites.
Neutron star structure and the neutron radius of Pb-208
Horowitz, C. J. and Piekarewicz, J · 2001
Earlier work this paper cites.
Smote: Synthetic minority over-sampling technique
Chawla, N. V., Bowyer, K. W., Hall, L. O., and Kegelmeyer, W. P · 2002
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Morin, F. and Bengio, Y · 2005
Earlier work this paper cites.
The limits of the nuclear landscape
Erler, J., Birge, N., Kortelainen, M., Nazarewicz, W., Olsen, E., Perhac, A., and Stoitsov, M. V · 2012
Earlier work this paper cites.
The maximum mass and radius of neutron stars and the nuclear symmetry energy
Gandolfi, S., Carlson, J., and Reddy, S · 2012
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
Bordes, A., Usunier, N., Garcia-Duran, A., Weston, J., and Yakhnenko, O · 2013
Earlier work this paper cites.
Surface diffuseness correction in global mass formula
Wang, N., Liu, M., Wu, X., and Meng, J · 2014
Earlier work this paper cites.
Learning entity and relation embeddings for knowledge graph completion
Lin, Y., Liu, Z., Sun, M., Liu, Y., and Zhu, X · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Schroff, F., Kalenichenko, D., and Philbin, J · 2015
Earlier work this paper cites.
Deep metric learning via lifted structured feature embedding
Oh Song, H., Xiang, Y., Jegelka, S., and Savarese, S · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class n-pair loss objective
Sohn, K · 2016
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Modeling relational data with graph convolutional networks
Schlichtkrull, M., Kipf, T. N., Bloem, P., Van Den Berg, R., Titov, I., and Welling, M · 2018
Cited alongside, same era.
Rotate: Knowledge graph embedding by relational rotation in complex space
Sun, Z., Deng, Z.-H., Nie, J.-Y., and Tang, J · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Mask-predict: Parallel decoding of conditional masked language models
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Later among the works it cites.
Multi-task learning on nuclear masses and separation energies with the kernel ridge regression
Wu, X., Lu, Y., and Zhao, P · 2022
Later among the works it cites.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J · 2022
Later among the works it cites.
Nuclear binding energies in artificial neural networks
Zeng, L.-X., Yin, Y.-Y., Dong, X.-X., and Geng, L.-S · 2022
Later among the works it cites.
Physics of language models: Part 3.2, knowledge manipulation
Allen-Zhu, Z. and Li, Y · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ghazvininejad, M., Levy, O., Liu, Y., and Zettlemoyer, L · 2019
Cited alongside, same era.
Learning attention-based embeddings for relation prediction in knowledge graphs
Nathani, D., Chauhan, J., Sharma, C., and Kaul, M · 2019
Cited alongside, same era.
Modeling tabular data using conditional gan
Xu, L., Skoularidou, M., Cuesta-Infante, A., and Veeramachaneni, K · 2019
Cited alongside, same era.
Methods for numeracy-preserving word embeddings
Sundararaman, D., Si, S., Subramanian, V., Wang, G., Hazarika, D., and Carin, L · 2020
Cited alongside, same era.
Structured denoising diffusion models in discrete state-spaces
Austin, J., Johnson, D. D., Ho, J., Tarlow, D., and Van Den Berg, R · 2021
Cited alongside, same era.
Machine learning the nuclear mass, 2021
Gao, Z., Wang, Y., Lü, H., Li, Q., Shen, C., and Liu, L · 2021
Cited alongside, same era.
Variational diffusion models
Kingma, D., Salimans, T., Poole, B., and Ho, J · 2021
Cited alongside, same era.
Closest in time.
The reversal curse: Llms trained on” a is b” fail to learn” b is a”
Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O · 2023
Closest in time.
Tabddpm: Modelling tabular data with diffusion models
Kotelnikov, A., Baranchuk, D., Rubachev, I., and Babenko, A · 2023
Closest in time.
Codi: Co-evolving contrastive diffusion models for mixed-type tabular synthesis
Lee, C., Kim, J., and Park, N · 2023
Closest in time.
Textbooks are all you need ii: phi-1.5 technical report
Li, Y., Bubeck, S., Eldan, R., Del Giorno, A., Gunasekar, S., and Lee, Y. T · 2023
Closest in time.
Learning to compress prompts with gist tokens
Mu, J., Li, X. L., and Goodman, N · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Pytorch fsdp: experiences on scaling fully sharded data parallel
Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., et al · 2023
Closest in time.
Physics of language models: Part 3.1, knowledge storage and extraction
Zhu, Z. A. and Li, Y · 2023
Closest in time.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T · 2024
Closest in time.