Skills
About the Role
Adobe’s Real-Time Customer Data Platform (RTCDP) helps power personalized experiences for leading global brands. As a Senior SRE on the RTCDP Datastores & AI/ML Ops team, you’ll ensure the platform remains reliable, scalable, and operationally excellent at worldwide scale.
This is a hands-on, high-ownership position focused on Day 2 production operations and core datastore engineering, with an expanding remit in operationalizing AI/ML services and workflows.
Responsibilities
- Drive day-to-day reliability for RTCDP services, including availability, performance, and durability aligned to SLOs
- Participate in on-call rotations and lead incident response, coordinating mitigation and recovery across SEV3–SEV1 events
- Conduct post-incident reviews and ensure follow-through on corrective actions
- Improve operational readiness through playbooks, runbooks, and continuous on-call process enhancements
- Partner with engineering teams to strengthen datastore and AI/ML operations as the service surface area grows
Requirements
- Strong experience in site reliability engineering or production operations for large-scale systems
- Proven ability to manage incidents and drive mitigation during high-severity events
- Hands-on background in datastore engineering and reliability fundamentals (performance, durability, availability)
- Experience building and improving operational tooling, playbooks, and operational processes
Benefits
- Opportunity to work on mission-critical customer data infrastructure at global scale
- High-ownership role with direct impact on reliability outcomes
- Collaborative environment focused on continuous improvement and operational excellence