**General information**
**Ref #** 22566
**Remote?** No
**Ally and Your Career**
*
Ally Financial only succeeds when its people do - and that's more than some cliché people put on job postings. We live this stuff! We see our people as, well, people - with interests, families, friends, dreams, and causes that are all important to them. Our focus is on the health and safety of our teammates as well as work-life balance and diversity and inclusion. From generous benefits to a variety of employee resource groups, we strive to build paths that encourage employees to stretch themselves professionally. We want to help you grow, develop, and learn new things. You're constantly evolving, so shouldn't your opportunities be, too?
**The Opportunity**
We are seeking a Director to lead Site Reliability Engineering (SRE) and Production Operations. This senior leader is accountable for the operational health, stability, resilience, and availability of production applications and platforms supporting Ally's Automotive and Insurance businesses. The role defines and executes the strategy for production operations and reliability engineering, driving continuous improvement and partnering across engineering, product, infrastructure, and business teams to deliver secure, stable, and scalable services.
At Ally, you get a startup feel, but experience the benefits of a company that's worked out the kinks and is fulfilling its purpose. We're always evolving and see that as a good thing. From owning our work to seeing its impact in the real world, our team is relentless in finding new ways technology can help make experiences better and help people. We are problem solvers, we value diverse thinking, we support one another, and we challenge ourselves to think bigger in the journey to deliver customer-obsessed tech solutions. To read more about what our tech team does, be sure to visit our tech blog at ally.tech
**The Work Itself**
Key Responsibilities
* Lead SRE strategy and production operations for critical application platforms, ensuring availability, resiliency, recoverability, and performance targets are consistently achieved.
* Own and evolve the operating model for production support, including incident, problem, and change risk management, as well as service restoration across the application portfolio.
* Drive adoption of SRE practices, including service level indicators (SLIs), service level objectives (SLOs), error budgets, operational readiness, and automation-first engineering approaches.
* Define the target-state SRE operating model and organization, including capacity planning, skill mix, and sourcing strategy (employees vs. contractors), to ensure sustainable 24x7 coverage aligned with business growth.
* Establish and institutionalize best practices across SRE and application sustainment, creating consistent, scalable standards for reliability engineering and operational execution.
* Establish and monitor operational health metrics, using data to identify systemic risks, improve reliability, reduce incident volume, and shorten recovery times.
* Provide executive leadership during major incidents, ensuring rapid coordination, clear communication, timely escalation, and durable corrective actions.
* Lead post-incident reviews and problem management efforts to resolve root causes, eliminate repeat issues, and strengthen operational discipline.
* Partner with product, engineering, infrastructure, and architecture teams to embed reliability, operability, and supportability into design, delivery, and release processes.
* Influence senior leaders across engineering, infrastructure, and business functions-including peer organizations and one level above-to align on reliability strategy, operating models, and investment priorities.
* Lead the evolution of traditional application sustainment toward a modern SRE-led model, ensuring a balanced transition that enhances reliability without disrupting critical support responsibilities.
* Drive automation and tooling investments that reduce manual effort, improve observability, streamline support processes, and increase engineering efficiency.
* Evaluate and quantify the impact of AI-driven operations (AIOps) and automation accelerators, driving data-informed adoption to improve reliability, efficiency, and cost outcomes.
* Define standards for monitoring, alerting, logging, capacity planning, and production readiness to strengthen proactive issue detection and service resilience.
* Influence cloud and platform transformation efforts by clarifying operational ownership, improving support models, and aligning reliability practices with modern engineering patterns.
* Establish a clear point of view on centralized versus distributed SRE models, shaping organizational design decisions that balance scale, accountability, and alignment with Agile delivery teams.
* Build, lead, and develop high-performing teams, fostering accountability, technical depth, and a culture of continuou