Skills
About the Role
Veeam is building a global Site Reliability Engineering (SRE) function to support the Veeam Data Cloud, our new SaaS platform. In this role, you’ll focus on our Government and Sovereign Cloud environment.
Because of clearance and access requirements, this team operates with restricted access to GOV infrastructure. You’ll help ensure the reliability, performance, and resilience of services in a highly regulated setting.
Responsibilities
- Own operational reliability for cloud services supporting Government and Sovereign Cloud customers
- Design, implement, and improve monitoring, alerting, and incident response practices
- Collaborate with engineering teams to drive SLO/SLI-based reliability improvements
- Contribute to automation efforts that reduce manual operations and improve system stability
- Participate in on-call and incident management activities as needed
Requirements
- Proven experience in Site Reliability Engineering or similar operations-focused roles
- Strong understanding of cloud infrastructure, reliability engineering, and troubleshooting
- Experience working in regulated environments or with compliance-driven operational requirements
- Ability to operate effectively with restricted access processes and security requirements
- Excellent communication skills and a collaborative mindset across cross-functional teams
Benefits
- Opportunity to work on mission-critical SaaS reliability for Government cloud environments
- Collaborate with a global team building a new SRE function
- Be part of an organization focused on data and AI trust, security, and resilience