LOG IN
SIGN UP
Canary Wharfian - Online Investment Banking & Finance Community.
Sign In
or continue with e-mail and password
Forgot password?
Don't have an account?
Join Canary Wharfian
or continue with e-mail and password
By signing up, you agree to our Terms & Conditions and Privacy Policy.

AI Data Engineer

ExperiencedNo visa sponsorship
Millennium logo

at Millennium

Hedge Funds

Posted 12 days ago

No clicks

**AI Data Engineer** at Millennium: Design, build, and maintain scalable ETL pipelines for diverse document sources, enhance retrieval quality via text extraction and enrichment, and ensure system resilience. Key skills: Python (4+ years), AI, LLMs, data pipeline (ETL), document processing (PDF, Office). Apply if you're a seasoned professional seeking to accelerate impact in a global investment firm.

Compensation
Not specified

Currency: Not specified

City
Not specified
Country
Not specified

Full Job Description

AI Data Engineer

About Millennium
Millennium is a global, diversified alternative investment firm, founded in 1989. Defined by evolution, innovation and focus, Millenniums mission is to deliver results for our investors.

Our people are empowered with both independence and support: the autonomy to pursue ideas with conviction and the backing of a global network committed to collaboration, disciplined risk management and continuous learning. With opportunities to deepen expertise and accelerate development, talent at Millennium is equipped to adapt, evolve and build lasting impact over time. Discover how transformative growth accelerates impact.

Our Israel office is located in the Bursa area of Ramat Gan.
This role is on-site.

As a global firm, proficiency in English is required.

Meet the Team
Core to the health and growth of our business, Millenniums Information Technology organization develops the flexible, scalable technology and advanced proprietary systems that support the firms multi-manager platform. The Core AI Development Team focuses on the engineering environment, data pipelines, and AI systems that help the firm apply Large Language Models in daily workflows, including enterprise retrieval and document intelligence capabilities.

What You'll Do
Design, build, and maintain scalable ETL and data ingestion pipelines that move documents from diverse source systems, including file shares, object stores, APIs, and databases, into the firms AI platform.
Develop robust document understanding workflows, including parsing, layout analysis, OCR, text extraction, metadata extraction, and normalization across heterogeneous formats such as PDF, Office documents, HTML, and images.
Implement chunking, cleaning, and enrichment strategies that improve retrieval quality and support downstream RAG systems.
Build change-detection, deduplication, and incremental update mechanisms to keep large document corpora synchronized efficiently and reliably.
Engineer pipelines for correctness, throughput, and resilience, with strong handling for malformed inputs, large files, and high-volume processing.
Establish data quality checks, observability, and metrics so ingestion issues are identified early and resolved quickly.
Partner with stakeholders to understand source systems and content requirements and translate them into reliable, production-ready ingestion solutions.
Stay current with advances in AI, LLMs, document AI, and retrieval techniques, and apply relevant improvements to the teams solutions.

What You Bring
4+ years of experience and strong proficiency in Python, including building data pipelines, services, and APIs.
Hands-on experience designing and developing ETL and data pipeline solutions, including processing large data volumes.
Experience with document processing and text extraction, including PDF and Office document parsing, OCR, and unstructured content handling.
Solid understanding of data modeling, transformation, and data quality best practices.
Experience designing, building, testing, and debugging high-performance, reliable systems.
Clear communication skills, with the ability to explain complex technical concepts to both technical and non-technical audiences.
Familiarity with RAG systems and the impact of ingestion on retrieval quality, including chunking strategies, embeddings, and vector stores, is a plus.

AI Data Engineer

Compensation

Not specified

City: Not specified

Country: Not specified

Millennium logo
Hedge Funds

12 days ago

No clicks

at Millennium

ExperiencedNo visa sponsorship

**AI Data Engineer** at Millennium: Design, build, and maintain scalable ETL pipelines for diverse document sources, enhance retrieval quality via text extraction and enrichment, and ensure system resilience. Key skills: Python (4+ years), AI, LLMs, data pipeline (ETL), document processing (PDF, Office). Apply if you're a seasoned professional seeking to accelerate impact in a global investment firm.

Full Job Description

AI Data Engineer

About Millennium
Millennium is a global, diversified alternative investment firm, founded in 1989. Defined by evolution, innovation and focus, Millenniums mission is to deliver results for our investors.

Our people are empowered with both independence and support: the autonomy to pursue ideas with conviction and the backing of a global network committed to collaboration, disciplined risk management and continuous learning. With opportunities to deepen expertise and accelerate development, talent at Millennium is equipped to adapt, evolve and build lasting impact over time. Discover how transformative growth accelerates impact.

Our Israel office is located in the Bursa area of Ramat Gan.
This role is on-site.

As a global firm, proficiency in English is required.

Meet the Team
Core to the health and growth of our business, Millenniums Information Technology organization develops the flexible, scalable technology and advanced proprietary systems that support the firms multi-manager platform. The Core AI Development Team focuses on the engineering environment, data pipelines, and AI systems that help the firm apply Large Language Models in daily workflows, including enterprise retrieval and document intelligence capabilities.

What You'll Do
Design, build, and maintain scalable ETL and data ingestion pipelines that move documents from diverse source systems, including file shares, object stores, APIs, and databases, into the firms AI platform.
Develop robust document understanding workflows, including parsing, layout analysis, OCR, text extraction, metadata extraction, and normalization across heterogeneous formats such as PDF, Office documents, HTML, and images.
Implement chunking, cleaning, and enrichment strategies that improve retrieval quality and support downstream RAG systems.
Build change-detection, deduplication, and incremental update mechanisms to keep large document corpora synchronized efficiently and reliably.
Engineer pipelines for correctness, throughput, and resilience, with strong handling for malformed inputs, large files, and high-volume processing.
Establish data quality checks, observability, and metrics so ingestion issues are identified early and resolved quickly.
Partner with stakeholders to understand source systems and content requirements and translate them into reliable, production-ready ingestion solutions.
Stay current with advances in AI, LLMs, document AI, and retrieval techniques, and apply relevant improvements to the teams solutions.

What You Bring
4+ years of experience and strong proficiency in Python, including building data pipelines, services, and APIs.
Hands-on experience designing and developing ETL and data pipeline solutions, including processing large data volumes.
Experience with document processing and text extraction, including PDF and Office document parsing, OCR, and unstructured content handling.
Solid understanding of data modeling, transformation, and data quality best practices.
Experience designing, building, testing, and debugging high-performance, reliable systems.
Clear communication skills, with the ability to explain complex technical concepts to both technical and non-technical audiences.
Familiarity with RAG systems and the impact of ingestion on retrieval quality, including chunking strategies, embeddings, and vector stores, is a plus.