LOG IN
SIGN UP
Canary Wharfian - Online Investment Banking & Finance Community.
Sign In
Forgot password?
Don't have an account?
or
Join Canary Wharfian
By signing up, you agree to our Terms & Conditions and Privacy Policy.
or

Site Reliability Engineer III

ExperiencedNo visa sponsorship
J.P. Morgan logo

at J.P. Morgan

Bulge Bracket Investment Banks

Posted 7 days ago

No clicks

**Site Reliability Engineer III** at JPMorgan Chase in Jersey City, NJ. Design, implement, and automate deployment approaches using CI/CD pipelines. Focus on application and infrastructure reliability, monitoring, and scalability. Collaborate with teams to drive operational improvements and contribute to AI-assisted incident response. Requires 3+ years of SRE experience, proficiency in scripting languages, and experience with monitoring tools (Grafana, Dynatrace, etc.). Familiarity with containerization (Docker, Kubernetes) and cloud environments (AWS/Azure) preferred.

Compensation
Not specified

Currency: Not specified

City
Not specified
Country
United States

Full Job Description

Location: Jersey City, NJ, United States

Theres nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. 

As a Site Reliability Engineer III at JPMorgan Chase within the Asset and Wealth Management team, you will solve complex and broad business problems with simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications and their associated infrastructure to independently decompose and iteratively improve on existing solutions. You are a significant contributor to your team by sharing your knowledge of end-to-end operations, availability, reliability, and scalability of your application or platform. 

Job responsibilities
  • Supports and collaborates with other engineers to design, develop, and implement deployment and reliability approaches using automated CI/CD pipelines
  • Implements infrastructure, configuration, and network as code for applications and platforms in your remit
  • Contributes to observability improvements including white and black box monitoring, service level objective alerting, and telemetry collection
  • Assists in identifying and resolving complex problems by using service level indicators and objectives to proactively address issues before they impact customers
  • Participates in incident response, triage, and post-incident analysis; helps document findings and remediation actions to prevent recurrence
  • Identifies opportunities to eliminate or automate remediation of recurring issues to reduce toil and improve overall operational stability
  • Uses enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements
  • Proactively recognizes roadblocks and identifies improvements to solve operational problems, including exploring new technologies where appropriate
  • Documents and shares knowledge within your organization via internal forums and communities of practice
  • Supports adoption of site reliability engineering best practices within your team
 
Required qualifications, capabilities, and skills
 
  • Formal training or certification on site reliability engineering concepts and 3+ years applied experience 
  • Foundational understanding of SRE culture and principles, including Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
  • General observability and monitoring understanding with working experience using industry-standard tooling (e.g., Grafana, Dynatrace, Prometheus, Datadog, Splunk, CloudWatch)
  • Proficiency in at least one scripting or programming language such as Python, Bash, or similar for tool development and operational support
  • Experience with incident and response management, including on-call participation and structured post-incident review
  • Knowledge of CI/CD pipelines and best practices using tools such as Jenkins, GitLab CI, or similar
  • Skills in automating repetitive tasks and managing configurations at scale using tools such as Ansible, Terraform, or similar
  • Familiarity with container technologies and container orchestration (e.g., Docker, Kubernetes)
  • Working knowledge of using enterprise-authorized AI capabilities to support SRE workflows, with strong validation habits and awareness of data sensitivity
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
 
Preferred qualifications, capabilities, and skills
 
  • Experience operating in cloud environments (AWS and/or Azure), including understanding of resiliency, scalability, and observability patterns
  • Familiarity with Kubernetes ingress, networking, and certificate deployment patterns
  • Experience improving infrastructure-as-code patterns (e.g., Terraform modules, reusable configurations)
  • Understanding of controls-focused operations in regulated environments, including change management discipline and audit support
  • Experience with service mesh, load balancing, and DNS troubleshooting (e.g., ALB/NLB, Route 53)
  • Drive to self-educate and evaluate emerging technologies in the SRE and cloud-native space
  • Strong communication skills with the ability to collaborate across different levels and stakeholder groups
Apply your skillsets to drive innovation and modernize the world's most complex and mission-critical systems

Site Reliability Engineer III

Compensation

Not specified

City: Not specified

Country: United States

J.P. Morgan logo
Bulge Bracket Investment Banks

7 days ago

No clicks

at J.P. Morgan

ExperiencedNo visa sponsorship

**Site Reliability Engineer III** at JPMorgan Chase in Jersey City, NJ. Design, implement, and automate deployment approaches using CI/CD pipelines. Focus on application and infrastructure reliability, monitoring, and scalability. Collaborate with teams to drive operational improvements and contribute to AI-assisted incident response. Requires 3+ years of SRE experience, proficiency in scripting languages, and experience with monitoring tools (Grafana, Dynatrace, etc.). Familiarity with containerization (Docker, Kubernetes) and cloud environments (AWS/Azure) preferred.

Full Job Description

Location: Jersey City, NJ, United States

Theres nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. 

As a Site Reliability Engineer III at JPMorgan Chase within the Asset and Wealth Management team, you will solve complex and broad business problems with simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications and their associated infrastructure to independently decompose and iteratively improve on existing solutions. You are a significant contributor to your team by sharing your knowledge of end-to-end operations, availability, reliability, and scalability of your application or platform. 

Job responsibilities
  • Supports and collaborates with other engineers to design, develop, and implement deployment and reliability approaches using automated CI/CD pipelines
  • Implements infrastructure, configuration, and network as code for applications and platforms in your remit
  • Contributes to observability improvements including white and black box monitoring, service level objective alerting, and telemetry collection
  • Assists in identifying and resolving complex problems by using service level indicators and objectives to proactively address issues before they impact customers
  • Participates in incident response, triage, and post-incident analysis; helps document findings and remediation actions to prevent recurrence
  • Identifies opportunities to eliminate or automate remediation of recurring issues to reduce toil and improve overall operational stability
  • Uses enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements
  • Proactively recognizes roadblocks and identifies improvements to solve operational problems, including exploring new technologies where appropriate
  • Documents and shares knowledge within your organization via internal forums and communities of practice
  • Supports adoption of site reliability engineering best practices within your team
 
Required qualifications, capabilities, and skills
 
  • Formal training or certification on site reliability engineering concepts and 3+ years applied experience 
  • Foundational understanding of SRE culture and principles, including Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
  • General observability and monitoring understanding with working experience using industry-standard tooling (e.g., Grafana, Dynatrace, Prometheus, Datadog, Splunk, CloudWatch)
  • Proficiency in at least one scripting or programming language such as Python, Bash, or similar for tool development and operational support
  • Experience with incident and response management, including on-call participation and structured post-incident review
  • Knowledge of CI/CD pipelines and best practices using tools such as Jenkins, GitLab CI, or similar
  • Skills in automating repetitive tasks and managing configurations at scale using tools such as Ansible, Terraform, or similar
  • Familiarity with container technologies and container orchestration (e.g., Docker, Kubernetes)
  • Working knowledge of using enterprise-authorized AI capabilities to support SRE workflows, with strong validation habits and awareness of data sensitivity
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
 
Preferred qualifications, capabilities, and skills
 
  • Experience operating in cloud environments (AWS and/or Azure), including understanding of resiliency, scalability, and observability patterns
  • Familiarity with Kubernetes ingress, networking, and certificate deployment patterns
  • Experience improving infrastructure-as-code patterns (e.g., Terraform modules, reusable configurations)
  • Understanding of controls-focused operations in regulated environments, including change management discipline and audit support
  • Experience with service mesh, load balancing, and DNS troubleshooting (e.g., ALB/NLB, Route 53)
  • Drive to self-educate and evaluate emerging technologies in the SRE and cloud-native space
  • Strong communication skills with the ability to collaborate across different levels and stakeholder groups
Apply your skillsets to drive innovation and modernize the world's most complex and mission-critical systems