Skills
About the Role
Join our Platform Engineering organization as a Staff Site Reliability Engineer supporting the Fabric team. You’ll help run and evolve the infrastructure that enables secure, reliable communication across systems and the public internet for our MongoDB products.
The Fabric team focuses on network architecture, service mesh, and edge load balancing, ensuring customer data is protected in transit while maintaining a globally connected multi-cloud environment.
Responsibilities
- Design, operate, and improve multi-cloud infrastructure that supports secure service-to-service and internet-facing communication
- Own reliability and operational excellence for networking components such as service mesh and edge load balancing
- Build and maintain observability and alerting signals to detect, diagnose, and prevent incidents
- Collaborate with engineering teams to harden deployment and operational workflows
- Drive continuous improvements to performance, availability, and security practices
Requirements
- Proven experience in Site Reliability Engineering or related infrastructure/operations roles
- Strong understanding of networking concepts and patterns for secure communication
- Hands-on experience with Kubernetes and modern service connectivity approaches (e.g., service mesh)
- Experience operating systems across multiple cloud environments
- Practical knowledge of observability/monitoring and incident response
- Ability to work effectively in a distributed team and support both remote and hybrid collaboration
Benefits
- Flexible work options: NYC HQ, Austin, Palo Alto, or San Francisco offices, or fully remote across North America
- Hybrid work accommodation when based in an office
- Opportunity to work on mission-critical infrastructure for a globally connected platform