Liftlab - Logo
Request a Demo
View all openings

Senior Site Reliability Engineer (Critical Incident Response)

Senior engineer on the P0 / critical incident team, responsible for leading rapid response to production-critical (P0/P1) incidents across LiftLab's multi-tenant platform. Works directly with the Critical Incident Manager to drive incidents to resolution, minimize client impact, and strengthen the reliability of the platform.

6-9 yrs expUSARemoteFull-timeReports to: Principal Data Engineer & Critical Incident Manager
  • Lead the technical response to P0/P1 production incidents, coordinating investigation and resolution across teams.
  • Make time-critical decisions (rollback vs. forward-fix) and drive incidents to resolution within SLA.
  • Own root-cause analysis and postmortems, and assign and track follow-up action items.
  • Build and maintain monitoring, alerting, and on-call processes to detect and prevent incidents.
  • Author and improve incident runbooks, escalation paths, and severity classification.
  • Provide US-hours coverage for rapid response and mentor junior incident engineers.

Express Your Interest

We review every profile personally. If your profile looks like a strong match, we'll reach out to you and move quickly.

By submitting, you agree to LiftLab's Privacy Policy. We'll only use your details to contact you about future opportunities.