LOG IN
SIGN UP
Canary Wharfian - Online Investment Banking & Finance Community.
Sign In
or continue with e-mail and password
Forgot password?
Don't have an account?
Join Canary Wharfian
or continue with e-mail and password
By signing up, you agree to our Terms & Conditions and Privacy Policy.

Senior Lead Infrastructure Engineer - Enterprise Technology Data Protection & Recovery

ExperiencedNo visa sponsorship
J.P. Morgan logo

at J.P. Morgan

Bulge Bracket Investment Banks

Posted 4 days ago

No clicks

**Senior Lead Infrastructure Engineer - Enterprise Technology Data Protection & Recovery** Lead cross-functional teams as a Senior Lead Infrastructure Engineer at JPMorgan Chase, owning the Resiliency Evidence Service (RES) product vision and backlog. Translate risk, regulatory needs into engineering guidance, partnering globally to drive firmwide recovery outcomes. Key responsibilities include: - Lead product ownership for RES, managing backlogs, stakeholder communications, and release milestones. - Provide hands-on infrastructure engineering leadership, defining reference architectures, evidence contracts, and integration patterns. - Steward RES end-to-end architecture, maintain architectural decision records, and manage availability SLOs. - Champion secure SDLC, threat modeling, and integration management; advance observability and SRE practices. - Represent RES to Risk, Controls, Audit, and Compliance partners, supporting critical environments and on-call obligations. Bring 5+ years in infrastructure engineering, proven product ownership experience, resiliency engineering expertise, and proficiency in Linux, containers (Kubernetes), Terraform, and AWS. Demonstrate strong programming language skills and data integration fluency.

Compensation
Not specified

Currency: Not specified

City
Not specified
Country
United States

Full Job Description

Location: OH, United States

Push the limits of whats possible with us as an experienced senior member of our product team. 

As a Senior Lead Infrastructure Engineer at JPMorgan Chase within the Enterprise Technology Data Protection & Recovery product line, you will operate at the intersection of infrastructure engineering, product ownership, and risk. You will own the product vision, roadmap, and backlog for the Resiliency Evidence Service (RES), the platform of record for capturing, validating, and reporting resiliency evidence across the firm, while providing hands-on technical leadership to the platform, application, and infrastructure teams that produce and consume that evidence. You will translate control objectives, regulatory expectations, and recovery targets (RTO/RPO) into concrete engineering guidance, evidence contracts, and delivery plans, then partner with global development and infrastructure teams to make those outcomes real.

Required qualifications, capabilities, and skills

  • Own the product vision, roadmap, and quarterly plan for the Resiliency Evidence Service; translate firmwide resiliency, risk, and audit objectives into a prioritized backlog with clear outcomes, success metrics, and release milestones.
  • Act as the single-threaded product owner: groom and prioritize the backlog, run sprint planning and reviews, manage cross-team dependencies, and communicate status, risks, and trade-offs to engineering, product, risk, and executive stakeholders.
  • Provide hands-on infrastructure engineering leadership: define reference architectures, evidence contracts (schemas, APIs, event formats), and integration patterns that partner teams use to emit resiliency evidence into RES.
  • Guide development and infrastructure teams on controls and control objectives from a risk perspective; interpret firmwide standards, recovery objectives, and audit findings into concrete engineering requirements, acceptance criteria, and evidence artifacts.
  • Partner with platform, application, and infrastructure teams to help them design, instrument, and produce the telemetry, attestations, restore proofs, and control evidence consumed by RES; run working sessions, review designs, and unblock adoption.
  •  Steward the end-to-end architecture of RES ingestion, storage, validation, and reporting surfaces; define domain boundaries, service contracts, and cross-service standards, and maintain architectural decision records.
  • Drive resilient service design for RES itself: manage availability SLOs, latency budgets, and error budgets; lead DR tests, restore validation exercises, and chaos/resilience drills for the platform and its dependencies.
  • Champion the secure SDLC across the ecosystem: threat modeling, SAST/DAST integration, dependency and SBOM management, secrets hygiene, encryption in transit and at rest, and robust authentication/authorization patterns for evidence pipelines.
  • Advance observability and SRE practices: instrument metrics, logs, and traces across ingestion and reporting paths; design actionable alerts; author runbooks; lead incident response and blameless postmortems for both RES and the evidence supply chain.
  • Support CI/CD and test quality across the product line: maintain an end-to-end understanding of how RES services, evidence pipelines, and partner integrations fit together and are tested (unit, integration, contract, performance); set expectations for coverage, quality gates, and progressive delivery (canary/blue-green) with rollback plans; review results, triage failures, and guide engineers to the likely root cause when tests or pipelines break.
  • Represent RES with Risk, Controls, Audit, and Compliance partners; produce data-backed status updates, escalation requests, and evidence packages that demonstrate control effectiveness and recovery posture.
  • Support critical environments (Development, QA, Simulation, Production) and lead management of on-call obligations for owned services, ensuring operational stability and timely resolution of incidents.
  • Contribute to the engineering community as an advocate of firmwide frameworks, tools, and best practices; add to team culture by fostering diversity, inclusion, and respect.

 

Required qualifications, capabilities, and skills

  • 5+ years experience in infrastructure, platform, or software engineering, with a proven track record leading cross-team delivery of production platforms in a large, regulated environment.
  • Demonstrated product ownership experience: authoring and maintaining a roadmap, running a prioritized backlog, leading sprint ceremonies, and managing stakeholder expectations across engineering, product, and risk audiences.
  • Deep knowledge of resiliency engineering and disaster recovery objectives with hands-on practice running DR tests, restore validation, and chaos/resilience exercises.
  • Strong understanding of technology risk and control frameworks: ability to interpret control objectives, regulatory expectations, and audit findings and translate them into concrete engineering requirements, evidence artifacts, and acceptance criteria.
  • Hands-on infrastructure engineering experience: Linux, containers and orchestration (Kubernetes), infrastructure as code (Terraform or equivalent), multi-AZ/region architectures, and cost-efficient scalability; hands-on experience with AWS and awareness of other cloud providers.
  • Working proficiency in at least one general-purpose programming language (e.g., Python, Java) sufficient to understand and design APIs, define integration tooling, prototype evidence pipelines, and review partner-team code.
  • Data and integration fluency: API design (REST/gRPC), event-driven patterns, schema evolution and compatibility, and practical experience with streaming or messaging platforms (Kafka or equivalent).
  • Observability discipline: instrumentation of metrics/logs/traces, alerting design, runbooks, incident management, and blameless postmortems.
  • Secure SDLC expertise: threat modeling, SAST/DAST integration, dependency risk/SBOM management, secrets handling, encryption standards, and robust authentication/authorization.
  • DevOps/SDLC: familiarity with build tools, CI/CD systems, and version control (Git) with automated pipeline governance.
  • Excellent written and verbal communication; able to move fluidly between deep technical review with engineers and outcome-oriented conversations with senior stakeholders.
  • Ability to independently tackle complex design, delivery, and prioritization problems with minimal oversight.

 

Preferred qualifications, capabilities, and skills

  • Prior ownership of a firmwide platform, control tool, or shared service consumed by many partner teams.
  • Experience defining evidence, control, or telemetry contracts (schemas, attestations, control mappings) and driving adoption across independent engineering teams.
  • Familiarity with data protection and recovery domains: backup and restore tooling, immutable storage, chain-of-custody controls, and auditability.
  • Experience with service mesh, API gateways, and contract governance at scale.
  • Knowledge of Windows/UNIX/Linux and shell scripting for operational tooling and automation.
  • Familiarity with modern frontend technologies for internal tooling or dashboards.
  • Experience establishing engineering standards across multiple teams and leading technical working groups.
  • Exposure to regulatory and audit regimes applicable to financial services resiliency (e.g., operational resilience, recovery and resolution planning).

 

Leadership and collaboration expectations

  • Act as a trusted technical and product leader across regions, influencing peers and decisionmakers to adopt leading-edge practices in resilience, risk, and platform engineering.
  • Set direction for partner teams by publishing clear evidence contracts, adoption guides, and roadmaps; then partner shoulder-to-shoulder with those teams to help them deliver.
  • Represent the product to senior stakeholders in Risk, Controls, Audit, and executive forums with data-backed narratives on progress, coverage, and residual risk.
  • Mentor engineers at multiple levels, elevate code quality and design rigor, and contribute to internal forums and tech talks to disseminate best practices.
  • Build strong working relationships with platform partners across Product Development, Infrastructure, Operations, Risk/Controls, and other firmwide functions to deliver competitive, scalable, and defensible solutions.
Lead Infrastructure Engineer and product owner driving the Resiliency Evidence Service and firmwide recovery outcomes.

Senior Lead Infrastructure Engineer - Enterprise Technology Data Protection & Recovery

Compensation

Not specified

City: Not specified

Country: United States

J.P. Morgan logo
Bulge Bracket Investment Banks

4 days ago

No clicks

at J.P. Morgan

ExperiencedNo visa sponsorship

**Senior Lead Infrastructure Engineer - Enterprise Technology Data Protection & Recovery** Lead cross-functional teams as a Senior Lead Infrastructure Engineer at JPMorgan Chase, owning the Resiliency Evidence Service (RES) product vision and backlog. Translate risk, regulatory needs into engineering guidance, partnering globally to drive firmwide recovery outcomes. Key responsibilities include: - Lead product ownership for RES, managing backlogs, stakeholder communications, and release milestones. - Provide hands-on infrastructure engineering leadership, defining reference architectures, evidence contracts, and integration patterns. - Steward RES end-to-end architecture, maintain architectural decision records, and manage availability SLOs. - Champion secure SDLC, threat modeling, and integration management; advance observability and SRE practices. - Represent RES to Risk, Controls, Audit, and Compliance partners, supporting critical environments and on-call obligations. Bring 5+ years in infrastructure engineering, proven product ownership experience, resiliency engineering expertise, and proficiency in Linux, containers (Kubernetes), Terraform, and AWS. Demonstrate strong programming language skills and data integration fluency.

Full Job Description

Location: OH, United States

Push the limits of whats possible with us as an experienced senior member of our product team. 

As a Senior Lead Infrastructure Engineer at JPMorgan Chase within the Enterprise Technology Data Protection & Recovery product line, you will operate at the intersection of infrastructure engineering, product ownership, and risk. You will own the product vision, roadmap, and backlog for the Resiliency Evidence Service (RES), the platform of record for capturing, validating, and reporting resiliency evidence across the firm, while providing hands-on technical leadership to the platform, application, and infrastructure teams that produce and consume that evidence. You will translate control objectives, regulatory expectations, and recovery targets (RTO/RPO) into concrete engineering guidance, evidence contracts, and delivery plans, then partner with global development and infrastructure teams to make those outcomes real.

Required qualifications, capabilities, and skills

  • Own the product vision, roadmap, and quarterly plan for the Resiliency Evidence Service; translate firmwide resiliency, risk, and audit objectives into a prioritized backlog with clear outcomes, success metrics, and release milestones.
  • Act as the single-threaded product owner: groom and prioritize the backlog, run sprint planning and reviews, manage cross-team dependencies, and communicate status, risks, and trade-offs to engineering, product, risk, and executive stakeholders.
  • Provide hands-on infrastructure engineering leadership: define reference architectures, evidence contracts (schemas, APIs, event formats), and integration patterns that partner teams use to emit resiliency evidence into RES.
  • Guide development and infrastructure teams on controls and control objectives from a risk perspective; interpret firmwide standards, recovery objectives, and audit findings into concrete engineering requirements, acceptance criteria, and evidence artifacts.
  • Partner with platform, application, and infrastructure teams to help them design, instrument, and produce the telemetry, attestations, restore proofs, and control evidence consumed by RES; run working sessions, review designs, and unblock adoption.
  •  Steward the end-to-end architecture of RES ingestion, storage, validation, and reporting surfaces; define domain boundaries, service contracts, and cross-service standards, and maintain architectural decision records.
  • Drive resilient service design for RES itself: manage availability SLOs, latency budgets, and error budgets; lead DR tests, restore validation exercises, and chaos/resilience drills for the platform and its dependencies.
  • Champion the secure SDLC across the ecosystem: threat modeling, SAST/DAST integration, dependency and SBOM management, secrets hygiene, encryption in transit and at rest, and robust authentication/authorization patterns for evidence pipelines.
  • Advance observability and SRE practices: instrument metrics, logs, and traces across ingestion and reporting paths; design actionable alerts; author runbooks; lead incident response and blameless postmortems for both RES and the evidence supply chain.
  • Support CI/CD and test quality across the product line: maintain an end-to-end understanding of how RES services, evidence pipelines, and partner integrations fit together and are tested (unit, integration, contract, performance); set expectations for coverage, quality gates, and progressive delivery (canary/blue-green) with rollback plans; review results, triage failures, and guide engineers to the likely root cause when tests or pipelines break.
  • Represent RES with Risk, Controls, Audit, and Compliance partners; produce data-backed status updates, escalation requests, and evidence packages that demonstrate control effectiveness and recovery posture.
  • Support critical environments (Development, QA, Simulation, Production) and lead management of on-call obligations for owned services, ensuring operational stability and timely resolution of incidents.
  • Contribute to the engineering community as an advocate of firmwide frameworks, tools, and best practices; add to team culture by fostering diversity, inclusion, and respect.

 

Required qualifications, capabilities, and skills

  • 5+ years experience in infrastructure, platform, or software engineering, with a proven track record leading cross-team delivery of production platforms in a large, regulated environment.
  • Demonstrated product ownership experience: authoring and maintaining a roadmap, running a prioritized backlog, leading sprint ceremonies, and managing stakeholder expectations across engineering, product, and risk audiences.
  • Deep knowledge of resiliency engineering and disaster recovery objectives with hands-on practice running DR tests, restore validation, and chaos/resilience exercises.
  • Strong understanding of technology risk and control frameworks: ability to interpret control objectives, regulatory expectations, and audit findings and translate them into concrete engineering requirements, evidence artifacts, and acceptance criteria.
  • Hands-on infrastructure engineering experience: Linux, containers and orchestration (Kubernetes), infrastructure as code (Terraform or equivalent), multi-AZ/region architectures, and cost-efficient scalability; hands-on experience with AWS and awareness of other cloud providers.
  • Working proficiency in at least one general-purpose programming language (e.g., Python, Java) sufficient to understand and design APIs, define integration tooling, prototype evidence pipelines, and review partner-team code.
  • Data and integration fluency: API design (REST/gRPC), event-driven patterns, schema evolution and compatibility, and practical experience with streaming or messaging platforms (Kafka or equivalent).
  • Observability discipline: instrumentation of metrics/logs/traces, alerting design, runbooks, incident management, and blameless postmortems.
  • Secure SDLC expertise: threat modeling, SAST/DAST integration, dependency risk/SBOM management, secrets handling, encryption standards, and robust authentication/authorization.
  • DevOps/SDLC: familiarity with build tools, CI/CD systems, and version control (Git) with automated pipeline governance.
  • Excellent written and verbal communication; able to move fluidly between deep technical review with engineers and outcome-oriented conversations with senior stakeholders.
  • Ability to independently tackle complex design, delivery, and prioritization problems with minimal oversight.

 

Preferred qualifications, capabilities, and skills

  • Prior ownership of a firmwide platform, control tool, or shared service consumed by many partner teams.
  • Experience defining evidence, control, or telemetry contracts (schemas, attestations, control mappings) and driving adoption across independent engineering teams.
  • Familiarity with data protection and recovery domains: backup and restore tooling, immutable storage, chain-of-custody controls, and auditability.
  • Experience with service mesh, API gateways, and contract governance at scale.
  • Knowledge of Windows/UNIX/Linux and shell scripting for operational tooling and automation.
  • Familiarity with modern frontend technologies for internal tooling or dashboards.
  • Experience establishing engineering standards across multiple teams and leading technical working groups.
  • Exposure to regulatory and audit regimes applicable to financial services resiliency (e.g., operational resilience, recovery and resolution planning).

 

Leadership and collaboration expectations

  • Act as a trusted technical and product leader across regions, influencing peers and decisionmakers to adopt leading-edge practices in resilience, risk, and platform engineering.
  • Set direction for partner teams by publishing clear evidence contracts, adoption guides, and roadmaps; then partner shoulder-to-shoulder with those teams to help them deliver.
  • Represent the product to senior stakeholders in Risk, Controls, Audit, and executive forums with data-backed narratives on progress, coverage, and residual risk.
  • Mentor engineers at multiple levels, elevate code quality and design rigor, and contribute to internal forums and tech talks to disseminate best practices.
  • Build strong working relationships with platform partners across Product Development, Infrastructure, Operations, Risk/Controls, and other firmwide functions to deliver competitive, scalable, and defensible solutions.
Lead Infrastructure Engineer and product owner driving the Resiliency Evidence Service and firmwide recovery outcomes.