LOG IN
SIGN UP
Canary Wharfian - Online Investment Banking & Finance Community.
Sign In
or continue with e-mail and password
Forgot password?
Don't have an account?
Join Canary Wharfian
or continue with e-mail and password
By signing up, you agree to our Terms & Conditions and Privacy Policy.

Cloud Operations - Service Reliability Engineer

ExperiencedNo visa sponsorship

Posted 7 days ago

No clicks

**Cloud Operations - Service Reliability Engineer** Engineer cloud-hosted services' reliability and resilience. Monitor, alert, automate using Bicep, Azure DevOps, GitHub, Ansible. Proactively identify issues, improve service health. Azure cloud engineering, Infrastructure as Code expertise essential. 4-5 years' experience, 2+ years in SRE/cloud ops role. ITIL, relevant accreditations preferred. Collaborate globally, improve services' supportability.

Compensation
Not specified

Currency: Not specified

City
Not specified
Country
United Kingdom

Full Job Description

Job description

What you will do

The Service Reliability Engineer is accountable for improving the reliability, observability and operational resilience of cloud-hosted services. The role focuses on monitoring, early issue identification, cloud engineering and automation, using tools and practices such as Bicep, Azure DevOps, GitHub and Ansible to support consistent, repeatable and well-governed service operation.
  • Monitoring, observability and alerting across cloud infrastructure, platform services and supported application environments;
  • Cloud engineering and automation, including Infrastructure as Code, deployment pipelines, configuration management and standards-led delivery; and
  • Promoting service resiliency through proactive issue identification, operational insight, automation and continuous improvement.
The role involves:
  • Supporting service reliability, observability and cloud engineering across the following areas:
    • Monitoring, observability and alerting for cloud-hosted services, including infrastructure health, service availability, performance signals and operational events Essential;
    • Azure public cloud engineering, including IaaS, PaaS, networking, identity, RBAC and platform diagnostics Essential;
    • Infrastructure as Code and automation using Bicep, Azure DevOps pipelines and GitHub-based source control and collaboration Essential;
    • Configuration management and standards automation using Ansible or equivalent tooling Preferred;
    • Experience of using or implementing monitoring solutions using Elastic Preferred;
    • Operational reporting, issue trend analysis and the development of actionable dashboards to support service improvement Preferred;
    • Experience of working across a broad range of systems, technologies and internal support teams Preferred.
  • Ensuring that monitoring and operational insight are effectively designed, implemented and understood so that services can be supported, improved and made more resilient.
  • Providing subject matter expertise in cloud operations, observability, automation and reliability engineering practices.
  • Working globally across cloud-hosted services and platform capabilities, independent of location.
  • Support the firms environmental goals and initiatives.

Monitoring, Reliability and Cloud Engineering

  • Works with internal technology teams to improve end-to-end observability for supported services, including:
    • Monitoring coverage for infrastructure, platform services and application components;
    • Actionable alerting that supports early identification of degradation, failure or operational risk;
    • Dashboards and reporting that help teams understand service health, trends and recurring issues;
    • Cloud engineering practices that use Bicep, Azure DevOps, GitHub and Ansible to deliver consistent and repeatable change; and
    • Operational standards that improve service resilience and reduce manual support effort.
  • Maintains appropriate documentation, including monitoring standards, known issues, operational patterns, troubleshooting guidance and support handbooks.

Service Delivery

  • Identify, diagnose and support resolution of incidents and problems by interpreting monitoring signals, operational telemetry and service behaviour.
  • Work with service-owning teams to improve the quality, relevance and routing of alerts so that operational issues can be detected and acted upon quickly.
  • Contribute to root cause analysis, problem management and continuous improvement activity by identifying recurring patterns, gaps in observability and opportunities for automation.

Build and Implementation

  • Provide specialist guidance to teams adopting cloud engineering patterns, Infrastructure as Code, deployment pipelines and automated configuration management.
  • Support implementation of monitoring and automation standards across new and existing services.
  • Ensure that operational documentation, handover materials and support guidance are created and are suitable for BAU operation.

Risk Management

  • Identify operational, reliability and supportability risks arising from gaps in monitoring, alerting, automation or cloud platform standards.
  • Refer to domain experts for guidance on specialised areas such as architecture, security, networking, database platforms and application design.
  • Participate in recovery, resilience and operational readiness activities to help prove that services can be supported effectively.

Quality, Methods & Tools

  • Strive for improvements to processes by promoting standardised patterns, automated controls, repeatable engineering practices and effective use of industry best practice.
  • Advocate for the use of source control, pipeline-based delivery, Infrastructure as Code and configuration management to improve quality, auditability and operational reliability.

What you will have

Business Competencies

  • Strong analytical and problem-solving skills, with a logical approach to issue identification, diagnosis and service improvement.
  • Technically curious, with an enthusiasm for understanding a broad set of systems, technologies and operational domains.
  • Ability to interpret monitoring data, identify patterns and translate operational insight into meaningful improvement activity.
  • Ability to make sound decisions under pressure and support effective incident response.
  • Strong commitment to service reliability, operational resilience and excellent customer service.
  • Commercial acumen, including an understanding of IT service costs, cloud consumption and how technology adds value to the business.
  • Ability to promote technical standards, automation and reliability practices using clear, business-friendly language.
  • Personal credibility; highly self-motivated self-starter who will undertake all activities to the highest professional standards.
  • Excellent communication skills, both orally and written.
  • Ability to operate within a wider team where there may be ambiguity and conflicting priorities.
  • Ability to build effective working relationships across a diverse set of internal teams and influence the adoption of monitoring, automation and cloud engineering standards.
  • Experience of working in a global environment across international locations with an appreciation of multiple cultures.

Knowledge

  • Practical knowledge of SRE principles, observability, incident response, problem management and operational resilience.
  • Detailed practical knowledge of Microsoft Azure infrastructure and platform services, including monitoring, diagnostics, RBAC, networking and automation.
  • Knowledge of Infrastructure as Code, source control and pipeline-based delivery using tools such as Bicep, Azure DevOps and GitHub.
  • Knowledge of configuration management and automation tooling such as Ansible.
  • Expected to develop a broad understanding of the technologies, systems and business working practices used by A&O Shearman.

Experience

  • Minimum 45 years IT experience with at least 2 years experience in a cloud operations, infrastructure, platform engineering, SRE or 3rd line support role.
  • Experience of monitoring, alerting, incident investigation and operational issue identification in a complex technology environment.
  • Experience of using or implementing monitoring using Elastic is desirable.
  • Experience using or supporting automation and delivery tooling such as Bicep, Azure DevOps, GitHub and Ansible.
  • Experience working with diverse internal teams to improve service supportability, resilience and operational standards.
  • Experience of working in an ITIL environment.

Qualifications

  • Ideally the candidate should have the following, or equivalent:
    • Minimum A level standard education or equivalent;
    • Accreditation in relevant technologies preferred; and
    • ITIL Foundation preferred.

Ideal candidate profile

The ideal candidate will be technically curious and motivated by understanding how different systems, platforms and operational processes fit together. They will enjoy working across a broad technology estate, using monitoring data and engineering insight to identify issues early, improve service resilience and guide teams towards standardised, automated and supportable ways of working.

NO AGENCIES PLEASE - A&O Shearman does not accept unsolicited CVs. For further information, please see our UK Recruitment Agency Policy and our commitment to direct sourcing here.
Job Code - Manager

Cloud Operations - Service Reliability Engineer

Compensation

Not specified

City: Not specified

Country: United Kingdom

A&O Shearman logo
Law

7 days ago

No clicks

at A&O Shearman

ExperiencedNo visa sponsorship

**Cloud Operations - Service Reliability Engineer** Engineer cloud-hosted services' reliability and resilience. Monitor, alert, automate using Bicep, Azure DevOps, GitHub, Ansible. Proactively identify issues, improve service health. Azure cloud engineering, Infrastructure as Code expertise essential. 4-5 years' experience, 2+ years in SRE/cloud ops role. ITIL, relevant accreditations preferred. Collaborate globally, improve services' supportability.

Full Job Description

Job description

What you will do

The Service Reliability Engineer is accountable for improving the reliability, observability and operational resilience of cloud-hosted services. The role focuses on monitoring, early issue identification, cloud engineering and automation, using tools and practices such as Bicep, Azure DevOps, GitHub and Ansible to support consistent, repeatable and well-governed service operation.
  • Monitoring, observability and alerting across cloud infrastructure, platform services and supported application environments;
  • Cloud engineering and automation, including Infrastructure as Code, deployment pipelines, configuration management and standards-led delivery; and
  • Promoting service resiliency through proactive issue identification, operational insight, automation and continuous improvement.
The role involves:
  • Supporting service reliability, observability and cloud engineering across the following areas:
    • Monitoring, observability and alerting for cloud-hosted services, including infrastructure health, service availability, performance signals and operational events Essential;
    • Azure public cloud engineering, including IaaS, PaaS, networking, identity, RBAC and platform diagnostics Essential;
    • Infrastructure as Code and automation using Bicep, Azure DevOps pipelines and GitHub-based source control and collaboration Essential;
    • Configuration management and standards automation using Ansible or equivalent tooling Preferred;
    • Experience of using or implementing monitoring solutions using Elastic Preferred;
    • Operational reporting, issue trend analysis and the development of actionable dashboards to support service improvement Preferred;
    • Experience of working across a broad range of systems, technologies and internal support teams Preferred.
  • Ensuring that monitoring and operational insight are effectively designed, implemented and understood so that services can be supported, improved and made more resilient.
  • Providing subject matter expertise in cloud operations, observability, automation and reliability engineering practices.
  • Working globally across cloud-hosted services and platform capabilities, independent of location.
  • Support the firms environmental goals and initiatives.

Monitoring, Reliability and Cloud Engineering

  • Works with internal technology teams to improve end-to-end observability for supported services, including:
    • Monitoring coverage for infrastructure, platform services and application components;
    • Actionable alerting that supports early identification of degradation, failure or operational risk;
    • Dashboards and reporting that help teams understand service health, trends and recurring issues;
    • Cloud engineering practices that use Bicep, Azure DevOps, GitHub and Ansible to deliver consistent and repeatable change; and
    • Operational standards that improve service resilience and reduce manual support effort.
  • Maintains appropriate documentation, including monitoring standards, known issues, operational patterns, troubleshooting guidance and support handbooks.

Service Delivery

  • Identify, diagnose and support resolution of incidents and problems by interpreting monitoring signals, operational telemetry and service behaviour.
  • Work with service-owning teams to improve the quality, relevance and routing of alerts so that operational issues can be detected and acted upon quickly.
  • Contribute to root cause analysis, problem management and continuous improvement activity by identifying recurring patterns, gaps in observability and opportunities for automation.

Build and Implementation

  • Provide specialist guidance to teams adopting cloud engineering patterns, Infrastructure as Code, deployment pipelines and automated configuration management.
  • Support implementation of monitoring and automation standards across new and existing services.
  • Ensure that operational documentation, handover materials and support guidance are created and are suitable for BAU operation.

Risk Management

  • Identify operational, reliability and supportability risks arising from gaps in monitoring, alerting, automation or cloud platform standards.
  • Refer to domain experts for guidance on specialised areas such as architecture, security, networking, database platforms and application design.
  • Participate in recovery, resilience and operational readiness activities to help prove that services can be supported effectively.

Quality, Methods & Tools

  • Strive for improvements to processes by promoting standardised patterns, automated controls, repeatable engineering practices and effective use of industry best practice.
  • Advocate for the use of source control, pipeline-based delivery, Infrastructure as Code and configuration management to improve quality, auditability and operational reliability.

What you will have

Business Competencies

  • Strong analytical and problem-solving skills, with a logical approach to issue identification, diagnosis and service improvement.
  • Technically curious, with an enthusiasm for understanding a broad set of systems, technologies and operational domains.
  • Ability to interpret monitoring data, identify patterns and translate operational insight into meaningful improvement activity.
  • Ability to make sound decisions under pressure and support effective incident response.
  • Strong commitment to service reliability, operational resilience and excellent customer service.
  • Commercial acumen, including an understanding of IT service costs, cloud consumption and how technology adds value to the business.
  • Ability to promote technical standards, automation and reliability practices using clear, business-friendly language.
  • Personal credibility; highly self-motivated self-starter who will undertake all activities to the highest professional standards.
  • Excellent communication skills, both orally and written.
  • Ability to operate within a wider team where there may be ambiguity and conflicting priorities.
  • Ability to build effective working relationships across a diverse set of internal teams and influence the adoption of monitoring, automation and cloud engineering standards.
  • Experience of working in a global environment across international locations with an appreciation of multiple cultures.

Knowledge

  • Practical knowledge of SRE principles, observability, incident response, problem management and operational resilience.
  • Detailed practical knowledge of Microsoft Azure infrastructure and platform services, including monitoring, diagnostics, RBAC, networking and automation.
  • Knowledge of Infrastructure as Code, source control and pipeline-based delivery using tools such as Bicep, Azure DevOps and GitHub.
  • Knowledge of configuration management and automation tooling such as Ansible.
  • Expected to develop a broad understanding of the technologies, systems and business working practices used by A&O Shearman.

Experience

  • Minimum 45 years IT experience with at least 2 years experience in a cloud operations, infrastructure, platform engineering, SRE or 3rd line support role.
  • Experience of monitoring, alerting, incident investigation and operational issue identification in a complex technology environment.
  • Experience of using or implementing monitoring using Elastic is desirable.
  • Experience using or supporting automation and delivery tooling such as Bicep, Azure DevOps, GitHub and Ansible.
  • Experience working with diverse internal teams to improve service supportability, resilience and operational standards.
  • Experience of working in an ITIL environment.

Qualifications

  • Ideally the candidate should have the following, or equivalent:
    • Minimum A level standard education or equivalent;
    • Accreditation in relevant technologies preferred; and
    • ITIL Foundation preferred.

Ideal candidate profile

The ideal candidate will be technically curious and motivated by understanding how different systems, platforms and operational processes fit together. They will enjoy working across a broad technology estate, using monitoring data and engineering insight to identify issues early, improve service resilience and guide teams towards standardised, automated and supportable ways of working.

NO AGENCIES PLEASE - A&O Shearman does not accept unsolicited CVs. For further information, please see our UK Recruitment Agency Policy and our commitment to direct sourcing here.
Job Code - Manager