Skills
About the Role
DAT is a next-generation SaaS technology company helping power transportation supply chain logistics with reliable, data-driven software used by millions of customers daily. We’re seeking a Staff Site Reliability Engineer to help ensure our services are highly available, resilient, and scalable as we continue to grow.
Responsibilities
- Design, build, and operate reliability-focused systems that improve service uptime, performance, and incident recovery
- Develop automation to reduce operational toil and strengthen observability across services
- Partner with engineering teams to implement SLOs/SLIs, monitoring, and continuous improvement practices
- Lead incident response efforts and drive post-incident learnings into durable technical fixes
- Improve deployment safety through reliable release processes, testing, and rollback strategies
Requirements
- Strong experience in SRE practices, reliability engineering, and production operations
- Proficiency with monitoring/observability tooling and building actionable dashboards/alerts
- Hands-on skills with cloud infrastructure and distributed systems
- Experience with automation and scripting to streamline operational workflows
- Excellent troubleshooting skills and the ability to collaborate effectively across teams
Benefits
- Opportunity to work on mission-critical systems used across North America
- Collaborative engineering culture focused on continuous improvement
- Career growth in a company investing in modern technology and infrastructure