2023

Eureka: Human-Level Reward Design via Coding Large Language Models

Ma, Yecheng Jason, Liang, William, Wang, Guanzhi et al.

Understand

Large Language Models (LLMs) have excelled as high-level semantic planners for sequential decision-making tasks.

  • However, harnessing them to learn complex low-level manipulation tasks, such as dexterous pen spinning, remains an open problem.
  • We bridge this fundamental gap and present Eureka, a human-level reward design algorithm powered by LLMs.
  • Eureka exploits the remarkable zero-shot generation, code-writing, and in-context improvement capabilities of state-of-the-art LLMs, such as GPT-4, to perform evolutionary optimization over reward code.

Reading the bibliography…