Senior Engineering Manager, Site Reliability at DC SCORES

Summary

Join Ditto's growing team as a Senior Engineering Manager of Site Reliability Engineering and lead a globally distributed SRE organization. You will develop engineering leaders, drive the adoption of SRE best practices, establish an incident management practice, and lead the architecture and execution of observability systems. Partner with other teams to build scalable, self-service reliability tooling and guide teams to define and implement SLIs, SLOs, and SLAs. Establish best-in-class documentation and model on-call excellence. Lead strategic programs to transform engineering culture toward reliability and design talent acquisition strategies. Play a central role in transforming Ditto’s engineering culture to prioritize reliability and resilience. This is a unique opportunity to shape enterprise-grade reliability, observability, and incident management for Ditto's systems.

Requirements

8+ years of experience in Site Reliability Engineering or related operational engineering roles
4+ years in engineering leadership, including managing other managers, with a track record of scaling high-performing teams
Demonstrated experience leading cultural transformation in engineering organizations, such as: Shifting from reactive firefighting to proactive reliability investment
Introducing SLOs and reducing incidents through continuous improvement programs
Empower and ultimately require teams to own their service health end-to-end
Expertise in cloud-native platforms (Kubernetes, Istio, etc.) and modern IaaS tooling (Terraform, Helm)
Strong background in observability and alerting strategy, using tools like Prometheus, Datadog, or OpenTelemetry
Hands-on experience with at least two major cloud providers (AWS, GCP, Azure)
Previous programing experience in one or more of the following languages (Go, Rust, Java, Python)
Excellent communicator and cross-functional partner, capable of influencing product and exec teams on engineering trade-offs
Adept at project and program management across multiple priorities, including balancing feature delivery with operational improvements
Exposure to chaos engineering, load testing, or resiliency modeling

Responsibilities

Lead and scale a globally distributed SRE organization, including managers and ICs, setting the long-term vision and execution plan for reliability at scale
Develop engineering leaders and senior talent, coaching on both technical depth and leadership maturity to create a high-trust, high-performance organization
Drive adoption of SRE best practices, including: Embedding SREs in product teams to influence design and early detection of failure modes
Defining production-readiness checklists and launch gates tied to SLOs
Championing error budgets as a shared accountability mechanism between product and reliability
Establish and evolve an incident management practice, including: Clear roles (Incident Commander, Scribe, Subject Matter Experts, CX & affected customer communication)
Blameless postmortems with systemic and meaningful remediations
Active tracking of incident themes and reliability KPIs, and reporting to senior leadership
Lead the architecture and execution of observability systems that offer real-time visibility into system health and customer experience
Partner with platform, infrastructure, and security teams to build scalable, self-service reliability tooling (e.g. circuit breakers, automated rollback, chaos testing frameworks)
Guide teams to define, implement, and iterate on SLIs, SLOs, and SLAs that are meaningful to end user experience
Establish best-in-class documentation and operational hygiene, including runbooks, architectural decision records (ADRs), and deep operational reviews
Model on-call excellence, including burnout prevention, clear handoffs, and leveraging automation and toil elimination
Lead strategic programs to transform engineering culture toward reliability, such as: Annual "Reliability Weeks", engineering health reviews
Incentivizing reliability work, such as inclusion in promotion criteria and roadmap planning
Designing systems to hold engineering teams accountable for the reliability of their respective systems
Design talent acquisition strategies, hiring criteria, and interview modules to build a team of exceptional talent
Design and implement a highly effective SRE org structure, including geo located teams, internal leadership and management lines, and integration/partnership points with other team
Play a central role in the transformation of Ditto’s engineering culture into a culture that prioritizes reliability and resilience of our mission critical software. Communicate & articulate this mission across the entire company in all hands, presentations, working sessions, and via enactment of strategic objectives

Preferred Qualifications

Experience building multi-tenant SaaS platforms with high uptime requirements
Familiarity with compliance-driven reliability environments (e.g., regulated industries)
Experience scaling DevOps platforms and internal tooling ecosystems
Passion for knowledge sharing through internal tech talks, RFCs, or public speaking

Benefits

Competitive salaries and meaningful equity
Health, dental, vision, life, and disability insurance
A 401(k) and flexible spending accounts
Private healthcare through Vitality
A pension plan
Flexible time off

Senior Engineering Manager, Site Reliability

DC SCORES

Summary

Requirements

Responsibilities

Preferred Qualifications

Benefits

Remote

DevOps

Manager

Share this job:

Similar Remote Jobs

Remote

DevOps

Manager

Remote

DevOps

Senior

Remote

DevOps

Senior

Remote

Software Development

Senior

Remote

Software Development

Senior

Remote

Software Development

Senior

Remote

DevOps

Senior

Trase

Remote

DevOps

Senior

Remote

DevOps

Senior