Member of Technical Staff, Evals

Posted

handshakeSan Francisco, CAfull_timesenior

Tech Stack

Responsibilities

  • Design and build evaluation frameworks, benchmarks, and methodologies for frontier LLMs, AI agents, multimodal models, and reinforcement-learning environments.
  • Develop reward models, programmatic verifiers, graders, and other feedback systems that make model behavior measurable and improvable.
  • Research what makes evaluations representative, difficult, reliable, and resistant to shortcutting or reward hacking.
  • Build systems for high-quality human data, including expert task design, annotation methodologies, data-quality signals, and data-attribution techniques.
  • Run fast, rigorous iteration loops: prototype, evaluate, interpret results, diagnose failure modes, and turn learnings into the next benchmark or system.

Soft Skills

Machine LearningArtificial IntelligenceLarge Language ModelsReinforcement LearningEvaluation FrameworksSoftware Engineering

Benefits

  • 401k
  • Commuter Benefits
  • Dental
  • Equity
  • Flexible PTO
  • Gym/Wellness
  • Health Insurance
  • Learning Budget
  • Mental Health
  • Parental Leave
  • Vision

Culture

High GrowthFast-PacedCollaborative Space

Requirements

Required: PhD in ML/AI, computer science, data science, or related fields (or equivalent research experience in industry)
Regions: Us

Get jobs like this in your inbox

Weekly Python hiring trends and salary data — free.

Join 8 engineers getting weekly insights

Get job market intelligence in your inbox

Free weekly insights on tech hiring trends, salaries, and in-demand stacks.

Already a subscriber? Sign in