
at J.P. Morgan
Bulge Bracket Investment BanksPosted 14 days ago
No clicks
**Senior Lead Site Reliability Engineer** at JPMorgan Chase in Jersey City, NJ seeks a seasoned professional to define strategy and manage a high-performing Production Management team supporting Audit and Credit Review. Lead business-critical Sales platform resiliency, drive incident management, and foster service maturity. Apply strong systems thinking, and proactively leverage AI for improved SRE workflows. Requirements include 10+ years in SRE, AI, and team leadership roles. Preferred skills span cloud services, automation, and modern platforms.
- Compensation
- Not specified
- City
- Not specified
- Country
- United States
Currency: Not specified
Full Job Description
Location: Jersey City, NJ, United States
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.
As a Lead Site Reliability Engineer at JPMorgan Chase within the
Lead the Production Management team supporting Audit and Credit Review, setting direction, priorities, and performance expectations.
Own stability, availability, resiliency, and end-to-end operational performance of business-critical Sales platforms, with clear accountability for outcomes.
Act as a senior escalation point during critical incidents, driving rapid triage, decisive coordination, and recovery actions aligned to business impact.
Drive operational consistency and service maturity through standardization, governance participation, service reviews, and disciplined support model integration for new capabilities.
Lead reliability improvements by applying systems thinking and root-cause practices, expanding SRE adoption (observability, monitoring, automation, operational analytics), and improving supportability with engineering teams.
Run strong incident, problem, and change managementmajor incident response, RCA and remediation to eliminate recurrence, and ensuring changes meet readiness standards (testing, monitoring, and resiliency/DR validation).
Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
Formal training or certification on site reliability engineering concepts and 10+ years applied experience
Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
Leadership experience across Production Support, Production Management, SRE, and Technology Operations teams, delivering stable and resilient production services.
Strong systems thinking and problem-solving capability to assess complex, cross-domain production issues and drive end-to-end resolution.
Demonstrated major incident leadership, coordinating effectively across teams to restore service rapidly and drive root-cause remediation.
Track record of partnering with Application Development, Product, Sales, and business stakeholders to improve reliability, service quality, and operational maturity.
Strong observability and service management expertise (Dynatrace, Splunk, Geneos, Grafana; ITIL Incident/Problem/Change/Availability), with excellent verbal/written communication and people leadership.
Demonstrated experience integrating support teams and standardizing operating models across multiple application groups to drive consistency and service maturity.
Hands-on technology experience across cloud (AWS/Azure/GCP), automation and scripting (Python, Shell, PowerShell, Ansible, Terraform), and modern distributed platforms (containers, microservices, Kubernetes/OpenShift).
SIMILAR OPPORTUNITIES

Senior Lead Site Reliability Engineer
J.P. Morgan
Added 15 days ago

Lead Site Reliability Engineer
HSBC
Added 15 days ago

Site Reliability Engineer / Senior Engineer
Deutsche Bank
Added 12 days ago

Junior Site Reliability Engineer
Capgemini
Added 12 days ago

Senior Service Reliability Engineer
Fitch Ratings
Added 7 days ago
