Tech Stack
Responsibilities
- Work across the inference stack to improve core performance metrics by diving deep into model execution.
- Identify bottlenecks and develop innovative optimizations.
- Collaborate closely with modeling and systems teams to experiment, measure, and ship improvements that meaningfully accelerate inference.
- Build expertise in advanced performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model execution strategies for MoE and large-scale architectures.
Benefits
- 401k
- Gym Membership
- Health Insurance
- Learning Budget
- Parental Leave
- Remote Stipend
- Remote Work
Culture
Cross-Functional TeamsMission-DrivenFast-PacedInclusive Hiring
Requirements
Regions: Us
Get jobs like this in your inbox
Weekly Git, Python, Rust hiring trends and salary data — free.
Join 8 engineers getting weekly insights
Get market intelligence in your inbox
Free weekly insights on tech hiring trends, salaries, and in-demand stacks.
Already a subscriber? Sign in
About cohere
Industry: saas
Size: medium
Cohere is the leading security-first enterprise AI company, building cutting-edge foundation AI models and end-to-end products designed to solve real-world business problems.
View company profile →Similar Jobs
Lead Member of Technical Staff, Inference Infrastructure
cohere · San Francisco
Senior Member of Technical Staff, Multimodal AI
cohere · San Francisco
Member of Technical Staff, Search
cohere · United States
Member of Technical Staff, MLE
cohere · San Francisco
Member of Technical Staff — Model Optimization and Inference (Experienced)
Nuance Communications · Seattle, Washington
$250k