Agent Post-Training, Frontier Evals and Environments Research

Posted

OpenAISan Franciscofull-time

Tech Stack

Responsibilities

  • Create ambitious RL environments to push models to their limits and measure frontier model capabilities, skills, and behaviors.
  • Develop new methodologies for automatically exploring the behavior of these models.
  • Dive deep into the science of measurement, including understanding scalability, reliability, and variance of evaluation methodology.
  • Help steer training for the largest training runs and see the future first.
  • Design scalable systems and processes to support continuous evaluation.

Culture

Mission-DrivenCustomer-ObsessedWork-Life Balance

Requirements

Regions: Us

Get jobs like this in your inbox

Weekly AWS, Go, Next.js hiring trends and salary data — free.

Join 8 engineers getting weekly insights

Get job market intelligence in your inbox

Free weekly insights on tech hiring trends, salaries, and in-demand stacks.

Already a subscriber? Sign in