Tech Stack
Responsibilities
- Propose and scope a new benchmark or evaluation technique in a domain APEX doesn’t yet cover well.
- Design task specifications and grading rubrics in partnership with Mercor’s network of vetted domain experts.
- Build and validate the benchmark by piloting tasks, calibrating scoring, and stress-testing for contamination.
- Run frontier models against your benchmark and analyze where and why they fail.
- Publish your results as a paper, an open dataset, a new leaderboard on APEX, or a methodology adopted internally.
Benefits
- Remote Work
Culture
Startup EnergyFast-PacedMentorship Program
Requirements
Preferred: Background in CS, ML, statistics, or an adjacent field
Regions: Us
Get jobs like this in your inbox
Weekly Next.js, TypeScript hiring trends and salary data — free.
Join 8 engineers getting weekly insights
Get job market intelligence in your inbox
Free weekly insights on tech hiring trends, salaries, and in-demand stacks.
Already a subscriber? Sign in
About mercor
Industry: saas
Size: small
Mercor is a leading AI data company valued at $10 billion, building the layer between human expertise and frontier models to organize human intelligence for the AI economy.
View company profile →Compensation
Base salary: $40,000 USD
Similar Jobs
Research Scientist, APEX Benchmarks
mercor · San Francisco
$15k
Base Labs Fellowship
baseten · San Francisco
$15k
Member of Technical Staff, Enterprise Evals Platform
mercor · San Francisco
$15k
Anthropic Fellows Program, AI Safety & Security
Anthropic · London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA
Remote
$15k
Anthropic Fellows Program, ML Systems & Reinforcement Learning
Anthropic · London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA
Remote
$15k