Skills
About the Role
Adobe’s Real-Time Customer Data Platform (RTCDP) enables personalized experiences for leading global brands. As a Senior Site Reliability Engineer (SRE), you’ll help keep RTCDP highly reliable, scalable, and operationally excellent at global scale—while expanding into the operational side of AI/ML services and workflows.
Responsibilities
- Own day-to-day reliability for RTCDP services, including availability, performance, and durability aligned to SLOs
- Participate in on-call rotations and lead incident response for SEV3–SEV1 events, driving mitigation and recovery
- Conduct post-incident reviews and drive follow-up improvements to prevent repeat issues
- Improve operational readiness through playbooks, runbooks, and on-call enhancements
Requirements
- Strong experience in production operations and reliability engineering
- Proven ability to manage incidents end-to-end, from triage through recovery and remediation
- Experience supporting large-scale datastore and platform services
- Familiarity with operationalizing AI/ML services or workflows is a plus
Benefits
- High-ownership role with meaningful impact on customer-facing systems
- Opportunity to grow into AI/ML operations alongside core datastore engineering
- Collaborative team environment focused on operational excellence