Skills
About the Role
Zscaler is seeking a Staff Site Reliability Engineer to strengthen the reliability and operational excellence of our production systems. You’ll help ensure our cloud services are resilient, performant, and secure, with an engineering approach grounded in measurable outcomes.
Responsibilities
- Design and operate production infrastructure with a focus on stability, scalability, and reliability
- Build and improve automation for deployment, monitoring, incident response, and recovery
- Partner with engineering teams to identify reliability risks and drive long-term fixes
- Develop and maintain observability practices, including metrics, alerting, and dashboards
- Lead post-incident reviews and implement actionable improvements to prevent recurrence
- Contribute to continuous improvement of operational standards and runbooks
Requirements
- Strong experience in SRE and production engineering practices
- Hands-on knowledge of Linux, networking fundamentals, and distributed systems
- Proficiency with scripting and automation (e.g., Python, Bash, or similar)
- Experience with monitoring/observability tooling and incident management
- Ability to collaborate across teams and communicate clearly during critical events
- Comfort working in fast-paced, production-focused environments
Benefits
- Opportunity to work on mission-critical, cloud-native services
- Collaborative culture that values constructive, honest debate
- Growth in an AI-forward organization focused on real-world impact