
at J.P. Morgan
Bulge Bracket Investment BanksPosted 6 days ago
No clicks
**Lead Infrastructure Cloud Engineer - DevOps & Python Focus** in Jersey City, NJ. As a Lead Engineer, **architect**, **deploy**, and **maintain** secure, scalable infrastructure. Proficient in **Kubernetes** (K8s) at the platform layer, **AWS** at scale, and **Terraform** for IAC. Drive **CI/CD**, **AI-assisted** software development, and SRE mindset for operational stability. Requires 5+ years of software engineering experience and proficient in the full SDLC. Preferred: IAM, **security** and **resiliency** depth, chaos engineering, and **cloud migration** experience.
- Compensation
- Not specified
- City
- Not specified
- Country
- United States
Currency: Not specified
Full Job Description
Location: Jersey City, NJ, United States
We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.
As a Lead Software Engineer at JPMorganChase within the Corporate Technology, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firms business objectives.
Job responsibilities
- Executes creative software solutions, design, development, and technical troubleshooting with ability to think beyond routine or conventional approaches to build solutions or break down technical problems
- Develops secure high-quality production code, and reviews and debugs code written by others
- Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
- Identifies opportunities to eliminate or automate remediation of recurring issues to improve overall operational stability of software applications and systems
- Leads evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture
- Leads communities of practice across Software Engineering to drive awareness and use of new and leading-edge technologies
- Adds to team culture of diversity, opportunity, inclusion, and respect
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and 5+ years applied experience
- Production Kubernetes depth cloud-managed and on-prem. Runs and troubleshoots K8s at the platform layer, not just "deploys to it": cluster add-ons and their failure modes (cert-manager, admission control, policy-as-code, GitOps, Helm, ingress & load-balancer controllers), RBAC, and control-plane, webhook issues.
- AWS at platform scale, multi-account. EKS, networking (VPC, load balancing, DNS), IAM, KMS secrets, and managed data stores. Fluent across resiliency patterns generally multi-region, active active, active passive, failover, and cell-based designs.
- Infrastructure-as-Code Terraform required. Solid hands-on Terraform: module-based, disciplined plan, apply, and safe change management.
- CI/CD & progressive delivery. End-to-end pipeline ownership blue, green and cell-based rollouts, GitOps, artifact build, and orchestration (e.g., Spinnaker, Jenkins) with strong safe-rollout and rollback judgment.
- Production operational stability, SRE mindset. Incident response, root-cause analysis, resiliency, observability-driven debugging, reducing toil, and hardening production.
- Some coding experience for internal tooling. Able to write scripts/automation and small internal tools (language-agnostic) to improve the platform not a deep hands-on IC requirement, just practical coding fluency.
- Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
- Proficient in all aspects of the Software Development Life Cycle
- IAM customer-authentication knowledge ForgeRock (AM, DS) or PingFederate, OAuth2, OIDC, SAML, PKI cert lifecycle. Helpful, but a strong platform engineer who ramps on our stack is fine.
- Resiliency chaos engineering (Game Days, fault injection).
- Secrets management (Vault, KMS, Secrets Manager). Observability depth (Splunk, Dynatrace).
- Managed database ops (Aurora, RDS, DB-change automation).
- On-prem cloud migration experience.
- "Platform-as-a-product" experience owning a platform other engineering teams depend on.
Certs as signal (not gates): AWS Solutions Architect, DevOps Pro, CKA, CKAD, ForgeRock, Ping.



