LOG IN
SIGN UP
Canary Wharfian - Online Investment Banking & Finance Community.
Sign In
Forgot password?
Don't have an account?
or
Join Canary Wharfian
By signing up, you agree to our Terms & Conditions and Privacy Policy.
or

Associate - AI Tooling Ops - Platform Reliability Engineer

ExperiencedNo visa sponsorship
Jefferies logo

at Jefferies

Investment Banking

Posted 5 days ago

No clicks

**Job Title:** Associate - AI Tooling Ops - Platform Reliability Engineer **Location:** Pune, India **Role Overview:** We are seeking a motivated Associate Platform Reliability Engineer (AI Tooling Ops) to join our global Platform Reliability Engineering team. This hands-on role focuses on ensuring the stability, reliability, scalability, and operational excellence of critical front-to-back business platforms supporting post-trade processing and operations. **Key Responsibilities:** 1. Partner with the team to design, build, and maintain AI Tooling Infrastructure running on AWS Kubernetes. 2. Monitor platform health, identify risks, and drive improvements in system reliability, performance, and availability. 3. Triage incidents, troubleshoot issues, communicate progress, and conduct post-incident reviews to minimize business impact. 4. Collaborate with stakeholders to design and implement scalable, resilient solutions. 5. Automate operational tasks to reduce manual support activities and enhance service efficiency. 6. Build and enhance deployment, monitoring, alerting, and observability capabilities across the platform stack. 7. Develop dashboards, alerts, and service health monitoring using Grafana, Datadog, Prometheus, and OpenTelemetry. 8. Analyze logs, metrics, and distributed traces to proactively identify system bottlenecks and reliability issues. 9. Support enterprise messaging and event-driven architectures, including Kafka-based platforms and integrations. 10. Troubleshoot complex issues across applications, middleware, databases, messaging platforms, infrastructure, and cloud environments. 11. Implement best practices for monitoring, observability, capacity planning, availability management, and operational excellence. 12. Participate in production support, problem management, release management, and change management activities. 13

Compensation
Not specified

Currency: Not specified

City
Not specified
Country
India

Full Job Description

Associate - AI Tooling Ops - Platform Reliability Engineer

Pune, India
.component-styling-wrapper-0 .intelligent-asset { flex-direction: column; align-items: center; }

Job Description

Associate Platform Reliability Engineer (AI Tooling Ops)

Location: Pune 

Role Overview

We are seeking a highly motivated Associate Platform Reliability Engineer (AI Tooling Ops) to join our global Platform Reliability Engineering team. This is a hands-on engineering role focused on ensuring the stability, reliability, scalability, and operational excellence of critical front-to-back business platforms supporting post-trade processing and operations.

The ideal candidate will have a strong software engineering foundation, production support experience, and a passion for automation, observability, and reliability engineering. You will work closely with development, infrastructure, and business teams to improve system resilience, enhance operational visibility, reduce manual intervention, and deliver highly available services.

 

Key Responsibilities

  • Partner with a high-performing global reliability engineering team to design, build, and maintain our AI Tooling Infrastructure running on AWS Kubernetes.
  • Monitor platform health, proactively identify risks, and drive improvements in system reliability, performance, and availability.
  • Perform incident triage, troubleshooting, communication, and post-incident reviews to minimize business impact and prevent recurrence.
  • Collaborate with engineering, infrastructure, and business stakeholders to design and implement scalable and resilient solutions.
  • Develop automation to reduce operational toil, eliminate manual support activities, and improve service efficiency.
  • Build and enhance deployment, monitoring, alerting, and observability capabilities across the platform stack.
  • Develop and maintain dashboards, alerts, and service health monitoring using Grafana, Datadog, Prometheus, and OpenTelemetry.
  • Analyze logs, metrics, and distributed traces to proactively identify system bottlenecks and reliability issues.
  • Support enterprise messaging and event-driven architectures, including Kafka-based platforms and integrations.
  • Troubleshoot complex issues across applications, middleware, databases, messaging platforms, infrastructure, and cloud environments.
  • Implement best practices for monitoring, observability, capacity planning, availability management, and operational excellence.
  • Participate in production support, problem management, release management, and change management activities.
  • Collaborate with regional teams across APAC, EMEA, and the Americas on strategic technology initiatives.
  • Drive continuous improvement initiatives focused on reliability, automation, monitoring, and operational efficiency.
 

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.
  • 3+ years of experience in Site Reliability Engineering (SRE), Platform Reliability Engineering (PRE), DevOps, Production Support, or Application Support.
  • Strong programming and scripting experience in one or more languages such as Python, Go, C#, Java, or C++.
  • Solid understanding of software engineering principles, data structures, algorithms, and system design.
  • Experience supporting and troubleshooting distributed applications in production environments.
  • Strong working knowledge of Linux/Unix and Windows Server environments.
  • Experience with relational and NoSQL databases, including performance analysis and troubleshooting.
  • Hands-on experience with observability and monitoring platforms including Grafana, Datadog, Prometheus, OpenTelemetry, Loki, and Jaeger.
  • Strong understanding of observability concepts, including Metrics, Logging, Tracing, Alerting, SLIs, SLOs, and Service Health Monitoring.
  • Experience creating operational dashboards, alerts, runbooks, and monitoring solutions to improve platform visibility.
  • Good understanding of event-driven architectures and enterprise messaging platforms such as Kafka.
  • Ability to troubleshoot message flows, APIs, middleware components, and distributed systems.
  • Understanding of incident management, problem management, root cause analysis, and operational support processes.
  • Familiarity with source control, CI/CD pipelines, Infrastructure as Code (IaC), and DevOps practices.
  • Excellent analytical and problem-solving skills with the ability to diagnose issues across the technology stack.
  • Strong verbal and written communication skills with the ability to engage both technical and business stakeholders.
  • Self-motivated, detail-oriented, and capable of working independently in a fast-paced environment.
 

Preferred Qualifications

Observability & Monitoring

  • Grafana
  • Prometheus
  • Datadog
  • OpenTelemetry
  • Loki
  • Jaeger

DevOps & Automation

  • Git
  • Jenkins
  • Ansible
  • Terraform
  • CI/CD Frameworks

Container & Platform Technologies

  • Docker
  • Kubernetes
  • OpenShift

Data & Messaging Platforms

  • Kafka
  • Redis
  • MongoDB
  • Elasticsearch

Cloud Technologies

  • AWS
  • Azure
  • GCP

 

About Us

Jefferies is a leading global, full-service investment banking and capital markets firm that provides advisory, sales and trading, research, and wealth and asset management services. With more than 40 offices around the world, we offer insights and expertise to investors, companies, and governments.

At Jefferies, we believe that diversity fosters creativity, innovation and thought leadership through the infusion of new ideas and perspectives. We have made a commitment to building a culture that provides opportunities for all employees regardless of our differences and supports a workforce that is reflective of the communities where we work and live. As a result, we are able to pool our collective insights and intelligence to provide fresh and innovative thinking for our clients.

Jefferies is an equal employment opportunity employer, and takes affirmative action to ensure that all qualified applicants will receive consideration for employment without regard to race, creed, color, national origin, ancestry, religion, gender, pregnancy, age, physical or mental disability, marital status, sexual orientation, gender identity or expression, veteran or military status, genetic information, reproductive health decisions, or any other factor protected by applicable law. We are committed to hiring the most qualified applicants and complying with all federal, state, and local equal employment opportunity laws. As part of this commitment, Jefferies will extend reasonable accommodations to individuals with disabilities, as required by applicable law.

.component-styling-wrapper-0 .apply-now-button, .apply-with-indeed-button { max-width: 270px; } .component-styling-wrapper-0 .apply-now-button, .apply-with-indeed-button { height: 60px; }
Apply Now

Job Info

  • Job Identification 4909
  • Job Category Information Technology
  • Posting Date 09/25/2026, 06:45 AM
  • Job Schedule Full time
  • Locations Gera Holdings Private Limited 200, Pune, 411001, IN

Similar Jobs

@media all and (min-width: 768px) { .component-styling-wrapper-0 .jobs-list__list, .component-styling-wrapper-0 .jobs-list__header, .component-styling-wrapper-0 .search-job-results__map-container { max-width: 100%; } }
  • openJobPreview(job.id), hasFocus: hasFocus " href="https://careers.jefferies.com/#en/sites/CX_1/job/4581">
    Associate - SDET - Platform Reliability Engineering
    Pune, India
    Posted on 07/20/2026
    Trending
  • openJobPreview(job.id), hasFocus: hasFocus " href="https://careers.jefferies.com/#en/sites/CX_1/job/4908">
    Associate - SRE - Platform Engineering
    Pune, India
    Posted on 09/25/2026
    Trending

Associate - AI Tooling Ops - Platform Reliability Engineer

Compensation

Not specified

City: Not specified

Country: India

Jefferies logo
Investment Banking

5 days ago

No clicks

at Jefferies

ExperiencedNo visa sponsorship

**Job Title:** Associate - AI Tooling Ops - Platform Reliability Engineer **Location:** Pune, India **Role Overview:** We are seeking a motivated Associate Platform Reliability Engineer (AI Tooling Ops) to join our global Platform Reliability Engineering team. This hands-on role focuses on ensuring the stability, reliability, scalability, and operational excellence of critical front-to-back business platforms supporting post-trade processing and operations. **Key Responsibilities:** 1. Partner with the team to design, build, and maintain AI Tooling Infrastructure running on AWS Kubernetes. 2. Monitor platform health, identify risks, and drive improvements in system reliability, performance, and availability. 3. Triage incidents, troubleshoot issues, communicate progress, and conduct post-incident reviews to minimize business impact. 4. Collaborate with stakeholders to design and implement scalable, resilient solutions. 5. Automate operational tasks to reduce manual support activities and enhance service efficiency. 6. Build and enhance deployment, monitoring, alerting, and observability capabilities across the platform stack. 7. Develop dashboards, alerts, and service health monitoring using Grafana, Datadog, Prometheus, and OpenTelemetry. 8. Analyze logs, metrics, and distributed traces to proactively identify system bottlenecks and reliability issues. 9. Support enterprise messaging and event-driven architectures, including Kafka-based platforms and integrations. 10. Troubleshoot complex issues across applications, middleware, databases, messaging platforms, infrastructure, and cloud environments. 11. Implement best practices for monitoring, observability, capacity planning, availability management, and operational excellence. 12. Participate in production support, problem management, release management, and change management activities. 13

Full Job Description

Associate - AI Tooling Ops - Platform Reliability Engineer

Pune, India
.component-styling-wrapper-0 .intelligent-asset { flex-direction: column; align-items: center; }

Job Description

Associate Platform Reliability Engineer (AI Tooling Ops)

Location: Pune 

Role Overview

We are seeking a highly motivated Associate Platform Reliability Engineer (AI Tooling Ops) to join our global Platform Reliability Engineering team. This is a hands-on engineering role focused on ensuring the stability, reliability, scalability, and operational excellence of critical front-to-back business platforms supporting post-trade processing and operations.

The ideal candidate will have a strong software engineering foundation, production support experience, and a passion for automation, observability, and reliability engineering. You will work closely with development, infrastructure, and business teams to improve system resilience, enhance operational visibility, reduce manual intervention, and deliver highly available services.

 

Key Responsibilities

  • Partner with a high-performing global reliability engineering team to design, build, and maintain our AI Tooling Infrastructure running on AWS Kubernetes.
  • Monitor platform health, proactively identify risks, and drive improvements in system reliability, performance, and availability.
  • Perform incident triage, troubleshooting, communication, and post-incident reviews to minimize business impact and prevent recurrence.
  • Collaborate with engineering, infrastructure, and business stakeholders to design and implement scalable and resilient solutions.
  • Develop automation to reduce operational toil, eliminate manual support activities, and improve service efficiency.
  • Build and enhance deployment, monitoring, alerting, and observability capabilities across the platform stack.
  • Develop and maintain dashboards, alerts, and service health monitoring using Grafana, Datadog, Prometheus, and OpenTelemetry.
  • Analyze logs, metrics, and distributed traces to proactively identify system bottlenecks and reliability issues.
  • Support enterprise messaging and event-driven architectures, including Kafka-based platforms and integrations.
  • Troubleshoot complex issues across applications, middleware, databases, messaging platforms, infrastructure, and cloud environments.
  • Implement best practices for monitoring, observability, capacity planning, availability management, and operational excellence.
  • Participate in production support, problem management, release management, and change management activities.
  • Collaborate with regional teams across APAC, EMEA, and the Americas on strategic technology initiatives.
  • Drive continuous improvement initiatives focused on reliability, automation, monitoring, and operational efficiency.
 

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.
  • 3+ years of experience in Site Reliability Engineering (SRE), Platform Reliability Engineering (PRE), DevOps, Production Support, or Application Support.
  • Strong programming and scripting experience in one or more languages such as Python, Go, C#, Java, or C++.
  • Solid understanding of software engineering principles, data structures, algorithms, and system design.
  • Experience supporting and troubleshooting distributed applications in production environments.
  • Strong working knowledge of Linux/Unix and Windows Server environments.
  • Experience with relational and NoSQL databases, including performance analysis and troubleshooting.
  • Hands-on experience with observability and monitoring platforms including Grafana, Datadog, Prometheus, OpenTelemetry, Loki, and Jaeger.
  • Strong understanding of observability concepts, including Metrics, Logging, Tracing, Alerting, SLIs, SLOs, and Service Health Monitoring.
  • Experience creating operational dashboards, alerts, runbooks, and monitoring solutions to improve platform visibility.
  • Good understanding of event-driven architectures and enterprise messaging platforms such as Kafka.
  • Ability to troubleshoot message flows, APIs, middleware components, and distributed systems.
  • Understanding of incident management, problem management, root cause analysis, and operational support processes.
  • Familiarity with source control, CI/CD pipelines, Infrastructure as Code (IaC), and DevOps practices.
  • Excellent analytical and problem-solving skills with the ability to diagnose issues across the technology stack.
  • Strong verbal and written communication skills with the ability to engage both technical and business stakeholders.
  • Self-motivated, detail-oriented, and capable of working independently in a fast-paced environment.
 

Preferred Qualifications

Observability & Monitoring

  • Grafana
  • Prometheus
  • Datadog
  • OpenTelemetry
  • Loki
  • Jaeger

DevOps & Automation

  • Git
  • Jenkins
  • Ansible
  • Terraform
  • CI/CD Frameworks

Container & Platform Technologies

  • Docker
  • Kubernetes
  • OpenShift

Data & Messaging Platforms

  • Kafka
  • Redis
  • MongoDB
  • Elasticsearch

Cloud Technologies

  • AWS
  • Azure
  • GCP

 

About Us

Jefferies is a leading global, full-service investment banking and capital markets firm that provides advisory, sales and trading, research, and wealth and asset management services. With more than 40 offices around the world, we offer insights and expertise to investors, companies, and governments.

At Jefferies, we believe that diversity fosters creativity, innovation and thought leadership through the infusion of new ideas and perspectives. We have made a commitment to building a culture that provides opportunities for all employees regardless of our differences and supports a workforce that is reflective of the communities where we work and live. As a result, we are able to pool our collective insights and intelligence to provide fresh and innovative thinking for our clients.

Jefferies is an equal employment opportunity employer, and takes affirmative action to ensure that all qualified applicants will receive consideration for employment without regard to race, creed, color, national origin, ancestry, religion, gender, pregnancy, age, physical or mental disability, marital status, sexual orientation, gender identity or expression, veteran or military status, genetic information, reproductive health decisions, or any other factor protected by applicable law. We are committed to hiring the most qualified applicants and complying with all federal, state, and local equal employment opportunity laws. As part of this commitment, Jefferies will extend reasonable accommodations to individuals with disabilities, as required by applicable law.

.component-styling-wrapper-0 .apply-now-button, .apply-with-indeed-button { max-width: 270px; } .component-styling-wrapper-0 .apply-now-button, .apply-with-indeed-button { height: 60px; }
Apply Now

Job Info

  • Job Identification 4909
  • Job Category Information Technology
  • Posting Date 09/25/2026, 06:45 AM
  • Job Schedule Full time
  • Locations Gera Holdings Private Limited 200, Pune, 411001, IN

Similar Jobs

@media all and (min-width: 768px) { .component-styling-wrapper-0 .jobs-list__list, .component-styling-wrapper-0 .jobs-list__header, .component-styling-wrapper-0 .search-job-results__map-container { max-width: 100%; } }
  • openJobPreview(job.id), hasFocus: hasFocus " href="https://careers.jefferies.com/#en/sites/CX_1/job/4581">
    Associate - SDET - Platform Reliability Engineering
    Pune, India
    Posted on 07/20/2026
    Trending
  • openJobPreview(job.id), hasFocus: hasFocus " href="https://careers.jefferies.com/#en/sites/CX_1/job/4908">
    Associate - SRE - Platform Engineering
    Pune, India
    Posted on 09/25/2026
    Trending