Production Services Specialist with Incident Management
Location: Chandler, AZ
Employment Type: Full Time
Responsibility
· The Production Services Specialist II provides advanced leadership across enterprise incident response, serving in Incident Commander and Major Incident Communications capacities for significant business-impacting events.
· The role orchestrates coordinated response efforts across technology domains, validates incident severity and business impact, drives time-sensitive decisions, and maintains clear communication with technical teams, business partners, operational leaders, and executive stakeholders.
· This specialist is accountable for establishing structure during critical events, maintaining situational awareness, aligning resources to the highest-priority recovery actions, and ensuring escalation paths are engaged appropriately.
· The position requires strong judgment under pressure, executive-level communication skills, broad understanding of production technology operations, and the ability to translate complex technical conditions into concise business-relevant updates.
· Lead end-to-end response activities for significant business-impacting incidents using disciplined major incident command practices, including clear role assignment, bridge control, time-boxed recovery checkpoints, decision logging, and structured restoration governance.
· Establish incident command structures and coordinate cross-functional response teams.
· Direct prioritization, escalation, resource allocation, containment, workaround, and restoration activities during active major incidents while ensuring technical teams remain aligned to the highest-probability recovery path.
· Maintain situational awareness across business impact, operational impact, risk, dependencies, and recovery progress.
· Facilitate executive, stakeholder, and major incident communications throughout the incident lifecycle, ensuring updates are timely, business-relevant, risk-aware, and consistent with established communication cadence and escalation protocols.
· Validate incident severity, business impact assessments, escalation decisions, and major incident engagement criteria.
· Drive rapid stabilization while minimizing customer, associate, and business disruption.
· Ensure compliance with incident management standards, governance requirements, operational procedures, audit expectations, severity models, escalation criteria, and post-incident documentation requirements.
· Partner with technical teams to identify recovery strategies, validate service restoration, preserve incident timelines, support root cause analysis activities, and translate post-incident findings into actionable improvement opportunities.
· Apply SRE-aligned concepts such as reliability engineering, alert hygiene, toil reduction, post-incident reviews, automation opportunities, service-level awareness, and continuous improvement to strengthen operational resiliency.