LOG IN
SIGN UP
Canary Wharfian - Online Investment Banking & Finance Community.
Sign In
Forgot password?
Don't have an account?
or
Join Canary Wharfian
By signing up, you agree to our Terms & Conditions and Privacy Policy.
or

Site Reliability Engineer III- Network

ExperiencedNo visa sponsorship
J.P. Morgan logo

at J.P. Morgan

Bulge Bracket Investment Banks

Posted 12 days ago

No clicks

**Site Reliability Engineer III - Network at JPMorgan Chase** Lead Network SRE, troubleshoot issues, automate operations, and improve observability. Key responsibilities involve guiding network problem management, driving automation, and collaborating with dev teams. Required skills include strong experience (3+ yrs) in SRE concepts, networking, and automation using Python, Shell, Ansible. SRE mindset, incident response, and cross-team coordination skills are essential. Leverage AI for risk identification and incident analysis. Preferred certifications: CCNP, experience with Cisco ACI. Work in a fast-growing tech field, modernizing critical systems.

Compensation
Not specified

Currency: Not specified

City
Not specified
Country
India

Full Job Description

Location: Hyderabad, Telangana, India

Theres nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. 

As a Site Reliability Engineer III at JPMorgan Chase within the Infrastructure Platforms team, you will solve complex and broad business problems with simple and straightforward solutions. Network SRE who owns troubleshooting and reliability improvements across network platforms. Leads problem management for recurring issues, drives automation-first operations, and partners with development teams to improve observability, alert quality, and resilience.



Job responsibilities

  • Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate, supporting adoption of site reliability engineering best practices within your team
  • Lead day-to-day operational ownership for network services, including complex troubleshooting and coordinated restoration.
  • Drive incident and problem management by running structured investigations, producing high-quality RCAs, and ensuring corrective/preventive actions are delivered.
  • Participate in major incident management, providing communications support, technical lead support, and mitigation execution.
  • Design and implement production-grade automation using Python, Shell, and Ansible (e.g., drift detection, change validation, automated diagnostics, safe rollout helpers).
  • Engineer and support software-defined networking capabilities, including SD-WAN, SDA, and broader SND.
  • Engineer and support routing and switching across enterprise networks.
  • Engineer and support security and L4L7 network components, including firewalls, load balancers, and proxies.
  • Improve reliability through standardization, guardrails, repeatable runbooks, continuous validation, and observability (dashboards, high-signal alerting, service health metrics) in partnership with developers/platform teams.
  • Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
 
 
 
Required qualifications, capabilities, and skills
 
  • Formal training or certification on site reliability engineering concepts and 3+ years applied experience 
  • Demonstrated experience in incident response and problem management, including end-to-end ownership of RCAs through closure.
  • Strong hands-on networking skills across enterprise routing/switching and security/L4L7 components.
  • Strong automation capability using Python, Shell, and Ansible in production operations.
  • SRE mindset with practical understanding of reliability concepts, NFRs, and risk analysis approaches (including familiarity with FMEA or equivalent methods).
  • Ability to work independently, prioritize effectively, and deliver with minimal oversight.
  • Experience supporting software-defined networking environments (e.g., SD-WAN, SDA, and related tooling).
  • Ability to build and operationalize monitoring/observability, including dashboards, alerting, and service health metrics.
  • Strong communication and coordination skills during high-severity incidents and cross-team restoration efforts.
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
 
 
Preferred qualifications, capabilities, and skills
 
  • Demonstrate experience with Cisco ACI / fabrics.
  • Hold relevant certifications such as CCNP (preferred), CCNA, or other vendor certifications.
  • Work effectively in a financial institution or other regulated environment.
 
 
Apply your skillsets to drive innovation and modernize the world's most complex and mission-critical systems

Site Reliability Engineer III- Network

Compensation

Not specified

City: Not specified

Country: India

J.P. Morgan logo
Bulge Bracket Investment Banks

12 days ago

No clicks

at J.P. Morgan

ExperiencedNo visa sponsorship

**Site Reliability Engineer III - Network at JPMorgan Chase** Lead Network SRE, troubleshoot issues, automate operations, and improve observability. Key responsibilities involve guiding network problem management, driving automation, and collaborating with dev teams. Required skills include strong experience (3+ yrs) in SRE concepts, networking, and automation using Python, Shell, Ansible. SRE mindset, incident response, and cross-team coordination skills are essential. Leverage AI for risk identification and incident analysis. Preferred certifications: CCNP, experience with Cisco ACI. Work in a fast-growing tech field, modernizing critical systems.

Full Job Description

Location: Hyderabad, Telangana, India

Theres nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. 

As a Site Reliability Engineer III at JPMorgan Chase within the Infrastructure Platforms team, you will solve complex and broad business problems with simple and straightforward solutions. Network SRE who owns troubleshooting and reliability improvements across network platforms. Leads problem management for recurring issues, drives automation-first operations, and partners with development teams to improve observability, alert quality, and resilience.



Job responsibilities

  • Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate, supporting adoption of site reliability engineering best practices within your team
  • Lead day-to-day operational ownership for network services, including complex troubleshooting and coordinated restoration.
  • Drive incident and problem management by running structured investigations, producing high-quality RCAs, and ensuring corrective/preventive actions are delivered.
  • Participate in major incident management, providing communications support, technical lead support, and mitigation execution.
  • Design and implement production-grade automation using Python, Shell, and Ansible (e.g., drift detection, change validation, automated diagnostics, safe rollout helpers).
  • Engineer and support software-defined networking capabilities, including SD-WAN, SDA, and broader SND.
  • Engineer and support routing and switching across enterprise networks.
  • Engineer and support security and L4L7 network components, including firewalls, load balancers, and proxies.
  • Improve reliability through standardization, guardrails, repeatable runbooks, continuous validation, and observability (dashboards, high-signal alerting, service health metrics) in partnership with developers/platform teams.
  • Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
 
 
 
Required qualifications, capabilities, and skills
 
  • Formal training or certification on site reliability engineering concepts and 3+ years applied experience 
  • Demonstrated experience in incident response and problem management, including end-to-end ownership of RCAs through closure.
  • Strong hands-on networking skills across enterprise routing/switching and security/L4L7 components.
  • Strong automation capability using Python, Shell, and Ansible in production operations.
  • SRE mindset with practical understanding of reliability concepts, NFRs, and risk analysis approaches (including familiarity with FMEA or equivalent methods).
  • Ability to work independently, prioritize effectively, and deliver with minimal oversight.
  • Experience supporting software-defined networking environments (e.g., SD-WAN, SDA, and related tooling).
  • Ability to build and operationalize monitoring/observability, including dashboards, alerting, and service health metrics.
  • Strong communication and coordination skills during high-severity incidents and cross-team restoration efforts.
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
 
 
Preferred qualifications, capabilities, and skills
 
  • Demonstrate experience with Cisco ACI / fabrics.
  • Hold relevant certifications such as CCNP (preferred), CCNA, or other vendor certifications.
  • Work effectively in a financial institution or other regulated environment.
 
 
Apply your skillsets to drive innovation and modernize the world's most complex and mission-critical systems