Tech Stack
Responsibilities
- Design and build evaluation frameworks, benchmarks, and methodologies for frontier LLMs, AI agents, multimodal models, and reinforcement-learning environments.
- Develop reward models, programmatic verifiers, graders, and other feedback systems that make model behavior measurable and improvable.
- Research what makes evaluations representative, difficult, reliable, and resistant to shortcutting or reward hacking.
- Build systems for high-quality human data, including expert task design, annotation methodologies, data-quality signals, and data-attribution techniques.
- Run fast, rigorous iteration loops: prototype, evaluate, interpret results, diagnose failure modes, and turn learnings into the next benchmark or system.
Soft Skills
Machine LearningArtificial IntelligenceLarge Language ModelsReinforcement LearningEvaluation FrameworksSoftware Engineering
Benefits
- 401k
- Commuter Benefits
- Dental
- Equity
- Flexible PTO
- Gym/Wellness
- Health Insurance
- Learning Budget
- Mental Health
- Parental Leave
- Vision
Culture
High GrowthFast-PacedCollaborative Space
Requirements
Required: PhD in ML/AI, computer science, data science, or related fields (or equivalent research experience in industry)
Regions: Us
Get jobs like this in your inbox
Weekly Python hiring trends and salary data — free.
Join 8 engineers getting weekly insights
Get job market intelligence in your inbox
Free weekly insights on tech hiring trends, salaries, and in-demand stacks.
Already a subscriber? Sign in
About handshake
Industry: saas
Size: large
Handshake organizes expert human knowledge to advance the AI economy, partnering with frontier labs on data, evaluation, and post-training challenges.
View company profile →Compensation
Equity: Equity in a fast-growing company
Similar Jobs
Member of Technical Staff, Post-Training
handshake · Remote (USA)
Remote
$2k
Member of Technical Staff, Data AI
handshake · San Francisco, CA
Member of Technical Staff - Evaluations
Reflection AI · San Francisco, CA
Member of Technical Staff, Data Flywheel
Reflection AI · New York, NY
Member of Technical Staff, Enterprise Evals Platform
mercor · San Francisco
$15k