Skills
About the Role
MongoDB’s Storage Layer Services (SLS) team is re-architecting the cloud storage layer that powers the next generation of MongoDB Atlas. As an SRE Manager for SLS, you’ll help deliver reliable, durable, and operationally safe distributed storage services that support multi-tenant customer workloads.
Responsibilities
- Partner with engineering teams to define SLOs and reliability targets for the storage layer
- Shape capacity planning and operational strategies to support growth and changing workload demands
- Ensure reliability, durability, and operational safety across core storage services
- Lead and scale a small team of senior SREs as founding members of the organization
- Drive execution against a multi-year roadmap for MongoDB cloud storage architecture
Requirements
- Experience leading Site Reliability Engineering (SRE) or similar reliability/operations functions
- Strong understanding of distributed systems reliability, durability, and incident management
- Proven ability to define and improve SLOs/SLIs and operational metrics
- Experience with capacity planning and performance/reliability tradeoffs in production environments
- Location in New York City (hybrid work model)
Benefits
- Opportunity to help build and lead an early team within a fast-growing organization
- Impactful role supporting critical infrastructure for MongoDB Atlas
- Hybrid working model based in New York City