Tech Stack
Responsibilities
- Design and build the ML eval stack at Exa, investigating how to evaluate search engines in an LLM world.
- Build scalable, reliable eval pipelines that track regressions, drift, and quality signals across billions of documents.
- Create golden datasets, synthetic benchmarks, agentic tasks, and real-world test suites that reflect how developers, agents, and humans actually use Exa.
- Partner closely with ML researchers, data engineers, infra engineers, and product to shape the feedback loops that improve search models.
Benefits
- Gym Membership
- Health Insurance
- Parental Leave
Culture
Mission-DrivenHigh GrowthCross-Functional Teams
Requirements
Regions: Us
Get jobs like this in your inbox
Weekly Express, Python, Rust hiring trends and salary data — free.
Join 8 engineers getting weekly insights
Get market intelligence in your inbox
Free weekly insights on tech hiring trends, salaries, and in-demand stacks.
Already a subscriber? Sign in
About exa
Industry: saas
Size: small
Exa is building a search engine for the AI era, powering agents, Fortune 500s, and AI labs with its Search API.
View company profile →Similar Jobs
Research, ML
exa · San Francisco, California
Member of Technical Staff (Data Scientist, Evals)
perplexity · San Francisco
Member of Technical Staff - Evaluations
Reflection AI · San Francisco, CA
Agent Post-Training, Frontier Evals and Environments Research
OpenAI · San Francisco
Developer Advocate
exa · San Francisco, California