Skills
About the Role
CoreWeave is seeking an Operations Engineering Manager to lead our Fleet Reliability Operations team. This group is central to how we deliver and maintain compute capacity—ensuring server nodes are provisioned, updated, and validated with speed and confidence.
Responsibilities
- Own operational reliability for the server fleet by guiding provisioning, updates, and triage workflows
- Lead execution of processes and tooling that configure and validate server nodes
- Drive continuous improvement across fleet operations to increase uptime, reduce risk, and improve maintainability
- Collaborate with cross-functional engineering teams to diagnose issues, prioritize fixes, and ensure consistent outcomes
- Oversee incident response and operational escalation paths related to fleet reliability
Requirements
- Experience leading operations or engineering teams in infrastructure, platforms, or large-scale environments
- Strong troubleshooting skills with a focus on reliability, observability, and operational rigor
- Proficiency with automation and tooling used to provision, configure, and validate systems
- Ability to manage complex operational workflows and communicate clearly across teams
Benefits
- Opportunity to build reliability at scale for an AI-focused infrastructure platform
- Work with a mission-driven team supporting critical compute capacity
- Competitive compensation and growth opportunities