
at J.P. Morgan
Bulge Bracket Investment BanksPosted 12 days ago
No clicks
**Software Engineer III, Site Reliability Engineering:** Leverage backend engineering and SRE principles in Tokyo-To, Japan, to enhance service availability and operational excellence. Key responsibilities include evolving backend services, automating CI/CD pipelines, implementing infrastructure-as-code, establishing observability, and collaborating across teams. Required skills: 3+ years of applicable SRE experience, proficiency in Python or Go, strong backend engineering, system design, and incident management. Preferred: production support, financial institution experience, cloud technologies, and front-end familiarity.
- Compensation
- Not specified
- City
- Tokyo
- Country
- Japan
Currency: Not specified
Full Job Description
Location: Tokyo-To, Japan
There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your engineering skills to drive innovation, stability, and resiliency across highly available platforms.
As a Software Engineer III Site Reliability Engineering (SRE) at JPMorganChase within Asset and Wealth Management, you will combine strong backend engineering fundamentals with SRE practices to improve end-to-end service availability, scalability, and operational excellence. You will design and deliver production-grade software, build automation to reduce toil, and strengthen reliability through well-defined observability, incident response, and continuous delivery practices.
As a seasoned member of an agile team, you will decompose ambiguous reliability challenges into clear, iterative improvementsdelivered through code and infrastructure-as-code.
Job Responsibilities
- Build and evolve backend services and reliability tooling with a focus on operational stability, resiliency, and maintainability.
- Design, implement, and improve reliability practices using automated continuous integration and continuous delivery (CI/CD) pipelines.
- Implement infrastructure, configuration, and network as code for the applications and platforms in your remit.
- Establish and operationalize observabilitywhite-box/black-box monitoring, telemetry, SLI/SLO definition, alerting strategies, and error budgetsand use these signals to prioritize reliability work.
- Collaborate with software engineers, stakeholders, and production support partners to resolve complex problems and proactively address issues before they impact customers.
- Participate in Major Incident Management (MIM): lead or support triage and stakeholder communications, coordinate cross-team resolution, and drive post-incident reviews that deliver measurable resiliency improvements.
- Use enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysisvalidating outputs and handling operational data per sensitivity and security requirements.
- Identify recurring toil and reliability risks, prioritize reusable solutions over one-off fixes, and track outcomes tied to SLOs and stability KPIs.
- Formal training or certification in software engineering/SRE concepts and 3+ years of applied experience.
- SRE mindset and working knowledge of reliability principles, including availability, scalability, incident management, and iterative improvement.
- Strong backend engineering experience and the ability to deliver secure, high-quality production code.
- Proficiency in Python or Go (preferred), or another modern language with willingness to ramp quickly.
- Hands-on experience with system design, application development, testing, and operational stability in a large-scale distributed environment.
- Observability experience, including telemetry collection, monitoring, and SLO-based alerting.
- Strong communication skills in English, with the ability to collaborate across teams and clearly document operational and technical decisions.
- Demonstrated ability to work effectively with enterprise-authorized AI-assisted engineering tools within the work environment as a core part of modern software and reliability engineering, including the ability to validate outputs for correctness, performance, and security, and to handle inputs/outputs in accordance with data sensitivity requirements.
- Production support / on-call experience, including incident triage, mitigation, root cause analysis, and driving post-incident improvements.
- Experience in a financial institution, preferably supporting Front Office platforms and time-sensitive, high-availability workloads.
- Exposure to cloud technologies and modern platform engineering practices.
- Familiarity with modern front-end technologies (not required, but beneficial depending on platform needs).



