Research Scientist - Human Alignment, Consumer Devices

Posted

OpenAISan Franciscofull-time

Tech Stack

Responsibilities

  • Develop RLHF and post-training methods for multimodal models.
  • Build reward models and preference-learning pipelines for adaptive, personalized model behavior.
  • Design datasets, rubrics, and evaluation frameworks that capture user preferences, contextual appropriateness, and long-term value in realistic tasks.
  • Run experiments on policy improvement using explicit feedback, implicit signals, and model-based grading.
  • Work on long-horizon evaluation problems, where model quality depends not just on a single response but on whether behavior improves outcomes over time.

Benefits

  • Health Insurance

Culture

Mission-DrivenCustomer-ObsessedCross-Functional TeamsInclusive Hiring

Requirements

Regions: Us

Get jobs like this in your inbox

Weekly AWS, Rust, TypeScript hiring trends and salary data — free.

Join 8 engineers getting weekly insights

Get market intelligence in your inbox

Free weekly insights on tech hiring trends, salaries, and in-demand stacks.

Already a subscriber? Sign in