Machine Learning Research Scientist, Evaluations
Posted
$165,600 USD
Tech Stack
Responsibilities
- Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents.
- Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities.
- Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them.
- Publish research findings in top-tier AI conferences.
Benefits
- 401k
- Equity
- Health Insurance
- Learning Budget
Culture
Inclusive HiringMission-Driven
Requirements
Required: Master's degree in Computer Science, Machine Learning, AI, or a related field
Preferred: Ph.D.
Regions: Us
Get jobs like this in your inbox
Weekly AWS, Next.js, TypeScript hiring trends and salary data — free.
Join 8 engineers getting weekly insights
Get job market intelligence in your inbox
Free weekly insights on tech hiring trends, salaries, and in-demand stacks.
Already a subscriber? Sign in
About Scale
Industry: saas
Size: medium
Scale provides high-quality data and full-stack technologies that power the world's leading generative AI models.
View company profile →Compensation
Base salary: $165,600 USD
Equity: Equity based compensation, subject to Board of Director approval
Similar Jobs
Machine Learning Research Scientist, Post-Training
Scale · San Francisco, CA; Seattle, WA; New York, NY
$166k
Research Scientist, Frontier Risk Evaluations
Scale · San Francisco, CA; New York, NY
$216k
Machine Learning Research Scientist Agents - Enterprise GenAI
Scale · San Francisco, CA; New York, NY
$265k
Research Scientist, Safety Post Training
Scale · San Francisco, CA; New York, NY
$216k
Research Scientist, Agent Robustness
Scale · San Francisco, CA; New York, NY
$216k