Skills
About the Role
Join our Site Reliability Engineering team, where you’ll help design and operate the global infrastructure that powers our services—especially the MongoDB Atlas platform. As customers expand worldwide, you’ll work on solutions that deliver low-latency performance and meet data sovereignty needs across regions. You’ll also help reduce operational overhead through infrastructure-as-code and self-healing systems, while improving visibility into system health.
Responsibilities
- Design, build, and evolve reliable infrastructure that supports global service delivery
- Partner closely with engineering teams to ensure smooth ownership boundaries and shared reliability goals
- Implement automation and infrastructure-as-code to lower operational burden
- Build self-healing mechanisms and improve monitoring to increase system observability
- Contribute to reliability strategies that support low-latency requests at scale
Requirements
- Experience building and operating production systems with a strong focus on reliability
- Proficiency with infrastructure-as-code practices
- Strong troubleshooting and operational mindset for complex distributed systems
- Ability to collaborate effectively across engineering teams
- Based in New York City and able to work a hybrid schedule
Benefits
- Hybrid work model based in New York City
- Opportunity to work on large-scale, globally distributed infrastructure
- Collaborative environment with engineering teams across the organization