LOG IN
SIGN UP
Canary Wharfian - Online Investment Banking & Finance Community.
Sign In
or continue with e-mail and password
Forgot password?
Don't have an account?
Join Canary Wharfian
or continue with e-mail and password
By signing up, you agree to our Terms & Conditions and Privacy Policy.

Lead Site Reliability Engineer

ExperiencedNo visa sponsorship
J.P. Morgan logo

at J.P. Morgan

Bulge Bracket Investment Banks

Posted 5 days ago

No clicks

**Lead Site Reliability Engineer (Plano, TX)** - Join JPMorgan Chase's Chief Technology Office as a Lead SRE driving IAM applications' reliability, security, and scalability.Balance speed, efficiency, and stability, influencing peers and senior stakeholders. Scale SRE adoption, set reliability expectations, and continuously improve with data-driven post-incident reviews. Lead AI-assisted reliability workflows, ensuring resiliency and security. Requires 5+ years in SRE, CI/CD tools, containers, and orchestration experience; demonstrated AI usage. Prefer data fluency, programming skills, and knowledge of modern service patterns.

Compensation
Not specified USD

Currency: $ (USD)

City
Plano
Country
United States

Full Job Description

Location: Plano, TX, United States

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

As a Lead Site Reliability Engineering at JPMorgan Chase within the Chief Technology Office, Identity & Access Management team, you are the non-functional requirement owner and champion for the applications in your remit. You are a key influencer in your teams strategic planning, driving continual improvement in customer experience, resiliency, security, scalability, monitoring, instrumentation, and automation of the software in your area. You act in a blameless, data-driven manner and navigate difficult situations with composure and tact.

Job responsibilities

  • Lead SRE practices that balance delivery speed, efficiency, and system stability 
  • Partner with engineering peers and senior stakeholders to drive strong, shared outcomes 
  • Scale SRE adoption across application and platform teams 
  • Set reliability expectations and show progress through stability and reliability metrics 
  • Run blameless, data-driven post-incident reviews and regular debriefs to turn lessons into improvements 
  • Build a continuous-improvement culture by gathering feedback and improving the customer experience 
  • Coach entry- to mid-level engineers and promote knowledge sharing through internal forums and communities 
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.

 

Required qualifications, capabilities, and skills

  • Formal training or certification in software engineering concepts plus 5+ years of applied experience 
  • Advanced knowledge of SRE principles and a track record of implementing SRE across application and platform teams while avoiding common pitfalls 
  • Experience leading technologists to manage and resolve complex technology issues at a firmwide level 
  • Ability to influence team culture by championing innovation and driving change 
  • Experience hiring, developing, and recognizing talent 
  • Hands-on experience with CI/CD tools (e.g., Jenkins, GitLab, Terraform) 
  • Experience with containers and orchestration (e.g., Docker, Kubernetes, ECS) 
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.

 

Preferred qualifications, capabilities, and skills

  • Ability to code, troubleshoot, and demonstrate strong data fluency
  • Strong troubleshooting skills across common networking technologies and issues 
  • Proficiency in at least one programming language (preferred: JavaScript, Go, Python
  • Working knowledge of modern service and integration patterns, including GraphQL fundamentalsevent-driven architecture (Kafka or equivalent), and observability/telemetry with OpenTelemetry
  • Strong understanding of OS platforms and distributed architecture

 

Strengthen critical identity systems to boost reliability, security, scalability, automation, and outstanding customer experiences.

Lead Site Reliability Engineer

Compensation

Not specified USD

City: Plano

Country: United States

J.P. Morgan logo
Bulge Bracket Investment Banks

5 days ago

No clicks

at J.P. Morgan

ExperiencedNo visa sponsorship

**Lead Site Reliability Engineer (Plano, TX)** - Join JPMorgan Chase's Chief Technology Office as a Lead SRE driving IAM applications' reliability, security, and scalability.Balance speed, efficiency, and stability, influencing peers and senior stakeholders. Scale SRE adoption, set reliability expectations, and continuously improve with data-driven post-incident reviews. Lead AI-assisted reliability workflows, ensuring resiliency and security. Requires 5+ years in SRE, CI/CD tools, containers, and orchestration experience; demonstrated AI usage. Prefer data fluency, programming skills, and knowledge of modern service patterns.

Full Job Description

Location: Plano, TX, United States

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

As a Lead Site Reliability Engineering at JPMorgan Chase within the Chief Technology Office, Identity & Access Management team, you are the non-functional requirement owner and champion for the applications in your remit. You are a key influencer in your teams strategic planning, driving continual improvement in customer experience, resiliency, security, scalability, monitoring, instrumentation, and automation of the software in your area. You act in a blameless, data-driven manner and navigate difficult situations with composure and tact.

Job responsibilities

  • Lead SRE practices that balance delivery speed, efficiency, and system stability 
  • Partner with engineering peers and senior stakeholders to drive strong, shared outcomes 
  • Scale SRE adoption across application and platform teams 
  • Set reliability expectations and show progress through stability and reliability metrics 
  • Run blameless, data-driven post-incident reviews and regular debriefs to turn lessons into improvements 
  • Build a continuous-improvement culture by gathering feedback and improving the customer experience 
  • Coach entry- to mid-level engineers and promote knowledge sharing through internal forums and communities 
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.

 

Required qualifications, capabilities, and skills

  • Formal training or certification in software engineering concepts plus 5+ years of applied experience 
  • Advanced knowledge of SRE principles and a track record of implementing SRE across application and platform teams while avoiding common pitfalls 
  • Experience leading technologists to manage and resolve complex technology issues at a firmwide level 
  • Ability to influence team culture by championing innovation and driving change 
  • Experience hiring, developing, and recognizing talent 
  • Hands-on experience with CI/CD tools (e.g., Jenkins, GitLab, Terraform) 
  • Experience with containers and orchestration (e.g., Docker, Kubernetes, ECS) 
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.

 

Preferred qualifications, capabilities, and skills

  • Ability to code, troubleshoot, and demonstrate strong data fluency
  • Strong troubleshooting skills across common networking technologies and issues 
  • Proficiency in at least one programming language (preferred: JavaScript, Go, Python
  • Working knowledge of modern service and integration patterns, including GraphQL fundamentalsevent-driven architecture (Kafka or equivalent), and observability/telemetry with OpenTelemetry
  • Strong understanding of OS platforms and distributed architecture

 

Strengthen critical identity systems to boost reliability, security, scalability, automation, and outstanding customer experiences.