
at J.P. Morgan
Bulge Bracket Investment BanksPosted 3 days ago
No clicks
**Senior Lead Software Engineer - Performance & Resiliency Engineering** Lead the optimization and stability enhancement of critical, high-volume platforms. Focus on latency reduction, throughput/TPS improvements, capacity planning, and peak-event readiness, while strengthening resiliency and release safety. Collaborate with application development and infrastructure teams to identify and fix stack bottlenecks, driving technical leadership and cross-team collaboration. Key responsibilities include load/stress testing, capacity modeling, production readiness, resiliency patterns, and safer releases via canary/progressive delivery. Proficiency in AWS, Kubernetes/EKS, ECS, and modern observability tools (Datadog, Dynatrace, etc.) is required. Candidate must possess a strong cloud/container background, proven experience in performance engineering, and the ability to lead-through-influence and mentor engineers.
- Compensation
- Not specified
- City
- Tampa
- Country
- United States
Currency: Not specified
Full Job Description
Location: Tampa, FL, United States
Performance and Resiliency Engineering Lead, Merchant Services (Senior Vice President or Executive Director)
Merchant Services is hiring an SVP/ED Performance & Resiliency Engineer to improve the performance and stability of critical, high-volume platforms. The role focuses on latency reduction, throughput/TPS improvements, capacity planning, and peak-event readiness, while also strengthening resiliency and release safety.
You will partner closely with application development teams and infrastructure/platform engineering to identify bottlenecks across the stack (application/runtime, database, network, compute, and platform), implement durable fixes, and raise engineering standards through technical leadership, mentorship, and strong cross-team collaboration.
Key responsibilities
- Lead performance engineering efforts: load/stress/soak testing, capacity modeling, performance tuning, and regression prevention (KPIs, guardrails, and acceptance criteria).
- Improve production readiness: observability (metrics/logs/traces/APM), actionable alerting, incident triage, and root-cause analysis leading to durable remediation.
- Strengthen resiliency patterns: timeouts/retries, circuit breakers, backpressure/rate limiting, graceful degradation, and failover readiness.
- Drive safer releases via canary/progressive delivery and automated rollback patterns.
- Optimize containerized workloads across EKS (primary) and ECS/other compute where applicable; drive autoscaling strategy and right-sizing.
Qualifications
- Senior experience in performance engineering for distributed systems and/or SRE-style reliability engineering in production.
- Strong cloud/container background (AWS + Kubernetes/EKS; ECS exposure beneficial).
- Experience with modern observability tooling (e.g., Datadog, Dynatrace, Grafana, OpenTelemetry, CloudWatch or equivalent).
- KEDA and/or Karpenter: large plus.
- Akamai: strongly preferred.
- Demonstrated ability to lead through influence, mentor engineers, and work effectively across teams.




