Fetching the paper…
Reading the bibliography…
We propose a novel approach for training large language models (LLMs) to adhere to objectives defined within a latent embedding space.
The equivalence of two extremum problems
Jack Kiefer and Jacob Wolfowitz · 1960
Earlier work this paper cites.
Optimal and efficient designs of experiments
Corwin L Atwood · 1969
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Collaborative filtering for implicit feedback datasets
Yifan Hu, Yehuda Koren, and Chris Volinsky · 2008
Earlier work this paper cites.
Geometric algorithms and combinatorial optimization , volume 2
Martin Grötschel, László Lovász, and Alexander Schrijver · 2012
Earlier work this paper cites.
Semi-supervised learning with deep generative models
Durk P Kingma, Shakir Mohamed, Danilo Jimenez Rezende, and Max Welling · 2014
Earlier work this paper cites.
The movielens datasets: History and context
F Maxwell Harper and Joseph A Konstan · 2015
Earlier work this paper cites.
Constraint-based question answering with knowledge graph
Junwei Bao, Nan Duan, Zhao Yan, Ming Zhou, and Tiejun Zhao · 2016
Earlier work this paper cites.
Enabling dark energy science with deep generative models of galaxy images
Siamak Ravanbakhsh, Francois Lanusse, Rachel Mandelbaum, Jeff Schneider, and Barnabas Poczos · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Improved variational autoencoders for text modeling using dilated convolutions
Zichao Yang, Zhiting Hu, Ruslan Salakhutdinov, and Taylor Berg-Kirkpatrick · 2017
Earlier work this paper cites.
Latent space oddity: on the curvature of deep generative models
Georgios Arvanitidis, Lars Kai Hansen, and Søren Hauberg · 2018
Earlier work this paper cites.
Automatic chemical design using a data-driven continuous representation of molecules
Rafael Gómez-Bombarelli, Jennifer N Wei, David Duvenaud, José Miguel Hernández-Lobato, Benjamín Sánchez-Lengeling, Dennis Sheberla, Jorge Aguilera-Iparraguirre, Timothy D Hirzel, Ryan P Adams, and Alán Aspuru-Guzik · 2018
Earlier work this paper cites.
Variational autoencoders for collaborative filtering
Dawen Liang, Rahul G Krishnan, Matthew D Hoffman, and Tony Jebara · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Fast approximate geodesics for deep generative models
Nutan Chen, Francesco Ferroni, Alexej Klushyn, Alexandros Paraschos, Justin Bayer, and Patrick van der Smagt · 2019
Earlier work this paper cites.
Minerl: A large-scale dataset of minecraft demonstrations
William H Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela Veloso, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Enabling hyperparameter optimization in sequential autoencoders for spiking neural data
Mohammad Reza Keshtkaran and Chethan Pandarinath · 2019
Earlier work this paper cites.
Justifying recommendations using distantly-labeled reviews and fine-grained aspects
Jianmo Ni, Jiacheng Li, and Julian McAuley · 2019
Cited alongside, same era.
Dr. vae: improving drug response prediction via modeling of drug perturbation effects
Ladislav Rampášek, Daniel Hidru, Petr Smirnov, Benjamin Haibe-Kains, and Anna Goldenberg · 2019
Cited alongside, same era.
The natural language of actions
Guy Tennenholtz and Shie Mannor · 2019
Cited alongside, same era.
Sampling-bias-corrected neural modeling for large corpus item recommendations
Xinyang Yi, Ji Yang, Lichan Hong, Derek Zhiyuan Cheng, Lukasz Heldt, Aditee Kumthekar, Zhe Zhao, Li Wei, and Ed Chi · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Cited alongside, same era.
Contextual bandits with large action spaces: Made practical
Yinglun Zhu, Dylan J Foster, John Langford, and Paul Mineiro · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Chain-of-verification reduces hallucination in large language models
Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston · 2023
Later among the works it cites.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Later among the works it cites.
Factual and personalized recommendations using language models and reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Contextual and sequential user embeddings for large-scale music recommendation
Casper Hansen, Christian Hansen, Lucas Maystre, Rishabh Mehrotra, Brian Brost, Federico Tomasi, and Mounia Lalmas · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Action space shaping in deep reinforcement learning
Anssi Kanervisto, Christian Scheller, and Ville Hautamäki · 2020
Cited alongside, same era.
Nvae: a deep hierarchical variational autoencoder
Arash Vahdat and Jan Kautz · 2020
Cited alongside, same era.
Mixed negative sampling for learning two-tower neural networks in recommendations
Ji Yang, Xinyang Yi, Derek Zhiyuan Cheng, Lichan Hong, Yang Li, Simon Xiaoming Wang, Taibai Xu, and Ed H Chi · 2020
Cited alongside, same era.
Geometrically enriched latent spaces
Georgios Arvanitidis, Soren Hauberg, and Bernhard Schölkopf · 2021
Cited alongside, same era.
Disease variant prediction with deep generative models of evolutionary data
Jonathan Frazer, Pascal Notin, Mafalda Dias, Aidan Gomez, Joseph K Min, Kelly Brock, Yarin Gal, and Debora S Marks · 2021
Cited alongside, same era.
Jihwan Jeong, Yinlam Chow, Guy Tennenholtz, Chih-Wei Hsu, Azamat Tulepbergenov, Mohammad Ghavamzadeh, and Craig Boutilier · 2023
Later among the works it cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi · 2023
Later among the works it cites.
Halueval: A large-scale hallucination evaluation benchmark for large language models
Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen · 2023
Later among the works it cites.
Representation-driven reinforcement learning
Ofir Nabati, Guy Tennenholtz, and Shie Mannor · 2023
Later among the works it cites.
An embedding approach for analyzing the evolution of research topics with a case study on computer science subdomains
Seyyed Reza Taher Harikandeh, Sadegh Aliakbary, and Soroush Taheri · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
Learning from prepandemic data to forecast viral escape
Nicole N Thadani, Sarah Gurev, Pascal Notin, Noor Youssef, Nathan J Rollins, Daniel Ritter, Chris Sander, Yarin Gal, and Debora S Marks · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Embedding in recommender systems: A survey
Xiangyu Zhao, Maolin Wang, Xinjian Zhao, Jiansheng Li, Shucheng Zhou, Dawei Yin, Qing Li, Jiliang Tang, and Ruocheng Guo · 2023
Later among the works it cites.
Recommender ecosystems: A mechanism design perspective on holistic modeling and optimization
Craig Boutilier, Martin Mladenov, and Guy Tennenholtz · 2024
Closest in time.
Detecting and preventing hallucinations in large vision language models
Anisha Gunjal, Jihan Yin, and Erhan Bas · 2024
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2024
Closest in time.
Demystifying embedding spaces using large language models
Guy Tennenholtz, Yinlam Chow, ChihWei Hsu, Jihwan Jeong, Lior Shani, Azamat Tulepbergenov, Deepak Ramachandran, Martin Mladenov, and Craig Boutilier · 2024
Closest in time.
Aligning language models with human preferences via a bayesian approach
Jiashuo Wang, Haozhao Wang, Shichao Sun, and Wenjie Li · 2024
Closest in time.