Skills
About the Role
Madison Reed is looking for a hands-on Senior SRE / AI Platform DevOps Engineer to build, operate, and scale the infrastructure powering our AI-driven services, agents, and orchestration platforms.
You’ll work at the intersection of site reliability engineering, cloud infrastructure, DevOps automation, observability, and AI operations—ensuring our systems are reliable, secure, scalable, and production-ready.
Responsibilities
- Design and maintain infrastructure and deployment practices for AI-enabled products
- Build reliable CI/CD workflows and automation for production releases
- Implement observability across telemetry, monitoring, alerting, and dashboards
- Own incident response practices, root-cause analysis, and reliability improvements
- Develop telemetry pipelines and monitoring frameworks for models, agents, and orchestration services
- Support governance processes for AI systems, including operational controls and best practices
Requirements
- Strong experience with SRE and/or DevOps in production environments
- Deep knowledge of cloud infrastructure and operational best practices
- Proven background in CI/CD, deployment automation, and release reliability
- Experience with monitoring/observability tooling and production incident management
- Automation mindset with a focus on cost-effective scaling
- Familiarity with AI operations concepts (model/agent/orchestration operationalization) is a plus
Benefits
- Work on infrastructure that directly enables AI-powered services
- High-impact role with ownership of reliability, security, and scalability
- Collaborate with engineering teams across platform, operations, and AI