LOG IN
SIGN UP
Canary Wharfian - Online Investment Banking & Finance Community.
Sign In
or continue with e-mail and password
Forgot password?
Don't have an account?
Join Canary Wharfian
or continue with e-mail and password
By signing up, you agree to our Terms & Conditions and Privacy Policy.

Platform Reliability Engineer

ExperiencedNo visa sponsorship
Millennium logo

at Millennium

Hedge Funds

Posted 7 days ago

No clicks

**Platform Reliability Engineer - Millennium:** Ensure Linux-based research & trading platform reliability. Manage incidents, identify risks, enhance monitoring, optimize systems, & automate processes. Requires 2+ yrs in SRE/DevOps, Linux expertise, storage tech knowledge, & familiarity with automation tools & HPC infrastructure. Join Millennium's collaborative infrastructure team supporting our global trading platform.

Compensation
Not specified

Currency: Not specified

City
Not specified
Country
Not specified

Full Job Description

Platform Reliability Engineer

About Millennium
Millennium is a global, diversified alternative investment firm, founded in 1989. Defined by evolution, innovation and focus, Millenniums mission is to deliver results for our investors.

Our people are empowered with both independence and support: the autonomy to pursue ideas with conviction and the backing of a global network committed to collaboration, disciplined risk management and continuous learning. With opportunities to deepen expertise and accelerate development, talent at Millennium is equipped to adapt, evolve and build lasting impact over time. Discover how transformative growth accelerates impact.

Meet the Team
Millenniums Infrastructure organization sits within Information Technology, a function that is core to the health and growth of the firm. The team designs, engineers, supports, and manages a robust server estate, systems virtualization, and core enterprise services that underpin a flexible, scalable technology environment. Working across a globally distributed platform, the team partners closely to improve reliability, reduce operational complexity, and support critical computing environments with a practical, collaborative approach.

What You'll Do

  • Help ensure the production reliability of the firms Linux-based research and trading platform as part of a globally distributed engineering team

  • Respond quickly and effectively to production infrastructure incidents to minimize disruption and restore service

  • Build a strong understanding of internal client needs and communicate priorities clearly to regional and global leadership

  • Identify operational risks, develop contingency plans, and implement solutions to strengthen platform resilience

  • Develop and enhance observability capabilities to monitor the health and performance of critical computing environments

  • Participate in a monthly on-call rotation and provide support to on-call engineers as needed

  • Improve operational efficiency through configuration management, systems optimization, and automation

  • Contribute to team knowledge through clear documentation, knowledge sharing, and maintainable code

What You Bring

  • 2+ years of experience in SRE, DevOps, or infrastructure engineering, ideally within the financial services industry

  • Strong knowledge of Linux systems internals, including kernel behavior, memory management, and performance optimization

  • Solid understanding of storage technologies, particularly in high-performance computing environments; GPFS experience is a plus

  • Broad infrastructure knowledge across networking and core services such as DNS, NTP/PTP, and NIS

  • Experience with automation, monitoring, and self-healing systems; Salt experience is a plus

  • Familiarity with container orchestration and virtualization technologies such as Kubernetes, Nomad, and VMware

  • Exposure to on-premises and cloud-based HPC infrastructure; operational knowledge of Slurm and GPU environments is a plus

  • Strong communication skills, a hands-on approach to problem-solving, and genuine curiosity about technology, automation, and AI-driven infrastructure improvements

Platform Reliability Engineer

Compensation

Not specified

City: Not specified

Country: Not specified

Millennium logo
Hedge Funds

7 days ago

No clicks

at Millennium

ExperiencedNo visa sponsorship

**Platform Reliability Engineer - Millennium:** Ensure Linux-based research & trading platform reliability. Manage incidents, identify risks, enhance monitoring, optimize systems, & automate processes. Requires 2+ yrs in SRE/DevOps, Linux expertise, storage tech knowledge, & familiarity with automation tools & HPC infrastructure. Join Millennium's collaborative infrastructure team supporting our global trading platform.

Full Job Description

Platform Reliability Engineer

About Millennium
Millennium is a global, diversified alternative investment firm, founded in 1989. Defined by evolution, innovation and focus, Millenniums mission is to deliver results for our investors.

Our people are empowered with both independence and support: the autonomy to pursue ideas with conviction and the backing of a global network committed to collaboration, disciplined risk management and continuous learning. With opportunities to deepen expertise and accelerate development, talent at Millennium is equipped to adapt, evolve and build lasting impact over time. Discover how transformative growth accelerates impact.

Meet the Team
Millenniums Infrastructure organization sits within Information Technology, a function that is core to the health and growth of the firm. The team designs, engineers, supports, and manages a robust server estate, systems virtualization, and core enterprise services that underpin a flexible, scalable technology environment. Working across a globally distributed platform, the team partners closely to improve reliability, reduce operational complexity, and support critical computing environments with a practical, collaborative approach.

What You'll Do

  • Help ensure the production reliability of the firms Linux-based research and trading platform as part of a globally distributed engineering team

  • Respond quickly and effectively to production infrastructure incidents to minimize disruption and restore service

  • Build a strong understanding of internal client needs and communicate priorities clearly to regional and global leadership

  • Identify operational risks, develop contingency plans, and implement solutions to strengthen platform resilience

  • Develop and enhance observability capabilities to monitor the health and performance of critical computing environments

  • Participate in a monthly on-call rotation and provide support to on-call engineers as needed

  • Improve operational efficiency through configuration management, systems optimization, and automation

  • Contribute to team knowledge through clear documentation, knowledge sharing, and maintainable code

What You Bring

  • 2+ years of experience in SRE, DevOps, or infrastructure engineering, ideally within the financial services industry

  • Strong knowledge of Linux systems internals, including kernel behavior, memory management, and performance optimization

  • Solid understanding of storage technologies, particularly in high-performance computing environments; GPFS experience is a plus

  • Broad infrastructure knowledge across networking and core services such as DNS, NTP/PTP, and NIS

  • Experience with automation, monitoring, and self-healing systems; Salt experience is a plus

  • Familiarity with container orchestration and virtualization technologies such as Kubernetes, Nomad, and VMware

  • Exposure to on-premises and cloud-based HPC infrastructure; operational knowledge of Slurm and GPU environments is a plus

  • Strong communication skills, a hands-on approach to problem-solving, and genuine curiosity about technology, automation, and AI-driven infrastructure improvements