Tech Stack
Responsibilities
- Evaluate user requests and AI model responses within the full relevant conversation context using complex customer policies.
- Distinguish between closely related labels, severity levels, and policy boundaries to choose defensible classifications.
- Write concise, evidence-based rationales that cite relevant policy language and conversation details.
- Identify policy gaps, contradictions, unclear definitions, and emerging edge cases and raise thoughtful questions.
- Participate actively in calibration discussions with evaluators, project leads, policy teams, and researchers.
Benefits
- Learning Budget
Culture
Fast-PacedStartup EnergyCross-Functional Teams
Requirements
Regions: Us
Get jobs like this in your inbox
Weekly Next.js, Rust, TypeScript hiring trends and salary data — free.
Join 8 engineers getting weekly insights
Get job market intelligence in your inbox
Free weekly insights on tech hiring trends, salaries, and in-demand stacks.
Already a subscriber? Sign in
About handshake
Industry: saas
Size: large
Handshake AI partners with leading AI research labs to make models safer and more robust through red teaming operations.
View company profile →Similar Jobs
AI Safety Policy Evaluator, Violence & Threats | Seattle Onsite
handshake · Seattle, WA
AI Red Teamer, LLM Generalist
handshake · Seattle, WA
AI Red Teamer (LLM Generalist)
handshake · Seattle, WA
Strategic Projects Lead, Safety
handshake · Seattle, WA
$2k
Software Engineer - AI Evaluator (Remote) - job post
handshake · San Jose, CA•Remote
Remote
Up to $208k