
at J.P. Morgan
Bulge Bracket Investment BanksPosted 12 days ago
No clicks
**Lead Software Engineer - Data & AI Platform Engineer** *Jersey City, NJ, USA* Lead Engineers drive software solutions, data pipeline design, and platform building. With 5+ years of software engineering experience, Python, Java, and SQL proficiency, you'll orchestrate Spark, Kafka, Flink, and Airflow. Cloud services like AWS S3, Glue, and Redshift are pluses. Leverage agile methodologies and collaborate with cross-functional teams to deliver data-driven outcomes. **Responsibilities:** - Design and build scalable data pipelines and ETL/ELT workflows - Develop data platforms, models, and catalogs with embedded governance and lineage - Translate data requirements into production-grade solutions - Review and debug code, and mentor junior engineers - Lead vendor evaluation sessions and architectural design discussions - Build Agentic Autonomous Lakehouse capability for automated pipeline provisioning - Foster inclusive team culture with diversity and respect **Requirements:** - Formal software engineering training (5+ years of applied experience) - Distributed data processing frameworks experience (Apache Spark, Flink) - Data modeling techniques knowledge (star schema, snowflake) and query optimization - AWS services proficiency (S3, Glue, Redshift) and EMR, Lake Formation, or equivalent
- Compensation
- Not specified USD
- City
- Jersey City
- Country
- United States
Currency: $ (USD)
Full Job Description
Location: Jersey City, NJ, United States
Executes creative software solutions, design, development, and technical troubleshooting with ability to think beyond routine or conventional approaches to build solutions or break down technical problems
Designs, builds, and maintains scalable data pipelines and ETL/ELT workflows for batch and real-time processing using Spark, Airflow, Kafka, and Flink
Develops data platform components including data cataloging, data quality frameworks, and semantic/metrics layers with embedded governance, lineage, and compliance standards
Implements data modeling strategies (fact and dimensional, wide tables) to support analytics, reporting, and downstream consumption
Partners with analytics teams, product managers, and business stakeholders to translate data requirements into production-grade solutions
Develops secure high-quality production code, and reviews and debugs code written by others
Identifies opportunities to eliminate or automate remediation of recurring issues to improve overall operational stability of software applications and systems
Leads evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture
Leads development of the Agentic Autonomous Lakehouse capability - automating governed self-service pipeline provisioning and lakehouse operations (health/cost/performance analysis, best-practice enforcement)
Leads communities of practice across Software Engineering to drive awareness and use of new and leading-edge technologies
Adds to team culture of diversity, opportunity, inclusion, and respect
Formal training or certification on software engineering concepts and 5+ years of applied experience
Hands-on practical experience delivering system design, application development, testing, and operational stability
Demonstrated professional experience focused on software engineering or data platform development
Advanced in one or more programming languages(s); Python, Java and SQL
Hands-on experience with distributed data processing frameworks such as Apache Spark and Flink
Solid understanding of data modeling techniques (star schema, snowflake) and query optimization
Experience designing and operating data pipelines on Databricks using orchestration tools such as Apache Airflow
Proficiency with cloud data services (AWS S3, Glue, Redshift, Athena, EMR, Lake Formation, or equivalent)
Experience engineering production-grade data platforms on Kubernetes with open catalog integration (e.g., Apache Iceberg, Unity Catalog, OpenMetadata) for scalable data discovery, lineage, and governance.
Advanced understanding of agile methodologies such as CI/CD, Application Resiliency, and Security
Experience developing Agentic AI, LLMs, RAG architectures, MCP, vector databases, and embedding-based retrieval systems
Hands-on familiarity with Data Platform and transformation framework development
Experience with data mesh or data product architectures
Proficiency with Infrastructure as Code (Terraform) and containerized deployments (Docker, Kubernetes)
Experience with data observability, quality, and metadata management tools
Experience with semantic layers, metrics stores, or BI platforms (Tableau, dbt Metrics)
SIMILAR OPPORTUNITIES
Lead Software Engineer - Data & AI Platform Engineer
Compensation
Not specified USD
City: Jersey City
Country: United States

**Lead Software Engineer - Data & AI Platform Engineer** *Jersey City, NJ, USA* Lead Engineers drive software solutions, data pipeline design, and platform building. With 5+ years of software engineering experience, Python, Java, and SQL proficiency, you'll orchestrate Spark, Kafka, Flink, and Airflow. Cloud services like AWS S3, Glue, and Redshift are pluses. Leverage agile methodologies and collaborate with cross-functional teams to deliver data-driven outcomes. **Responsibilities:** - Design and build scalable data pipelines and ETL/ELT workflows - Develop data platforms, models, and catalogs with embedded governance and lineage - Translate data requirements into production-grade solutions - Review and debug code, and mentor junior engineers - Lead vendor evaluation sessions and architectural design discussions - Build Agentic Autonomous Lakehouse capability for automated pipeline provisioning - Foster inclusive team culture with diversity and respect **Requirements:** - Formal software engineering training (5+ years of applied experience) - Distributed data processing frameworks experience (Apache Spark, Flink) - Data modeling techniques knowledge (star schema, snowflake) and query optimization - AWS services proficiency (S3, Glue, Redshift) and EMR, Lake Formation, or equivalent
Full Job Description
Location: Jersey City, NJ, United States
Executes creative software solutions, design, development, and technical troubleshooting with ability to think beyond routine or conventional approaches to build solutions or break down technical problems
Designs, builds, and maintains scalable data pipelines and ETL/ELT workflows for batch and real-time processing using Spark, Airflow, Kafka, and Flink
Develops data platform components including data cataloging, data quality frameworks, and semantic/metrics layers with embedded governance, lineage, and compliance standards
Implements data modeling strategies (fact and dimensional, wide tables) to support analytics, reporting, and downstream consumption
Partners with analytics teams, product managers, and business stakeholders to translate data requirements into production-grade solutions
Develops secure high-quality production code, and reviews and debugs code written by others
Identifies opportunities to eliminate or automate remediation of recurring issues to improve overall operational stability of software applications and systems
Leads evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture
Leads development of the Agentic Autonomous Lakehouse capability - automating governed self-service pipeline provisioning and lakehouse operations (health/cost/performance analysis, best-practice enforcement)
Leads communities of practice across Software Engineering to drive awareness and use of new and leading-edge technologies
Adds to team culture of diversity, opportunity, inclusion, and respect
Formal training or certification on software engineering concepts and 5+ years of applied experience
Hands-on practical experience delivering system design, application development, testing, and operational stability
Demonstrated professional experience focused on software engineering or data platform development
Advanced in one or more programming languages(s); Python, Java and SQL
Hands-on experience with distributed data processing frameworks such as Apache Spark and Flink
Solid understanding of data modeling techniques (star schema, snowflake) and query optimization
Experience designing and operating data pipelines on Databricks using orchestration tools such as Apache Airflow
Proficiency with cloud data services (AWS S3, Glue, Redshift, Athena, EMR, Lake Formation, or equivalent)
Experience engineering production-grade data platforms on Kubernetes with open catalog integration (e.g., Apache Iceberg, Unity Catalog, OpenMetadata) for scalable data discovery, lineage, and governance.
Advanced understanding of agile methodologies such as CI/CD, Application Resiliency, and Security
Experience developing Agentic AI, LLMs, RAG architectures, MCP, vector databases, and embedding-based retrieval systems
Hands-on familiarity with Data Platform and transformation framework development
Experience with data mesh or data product architectures
Proficiency with Infrastructure as Code (Terraform) and containerized deployments (Docker, Kubernetes)
Experience with data observability, quality, and metadata management tools
Experience with semantic layers, metrics stores, or BI platforms (Tableau, dbt Metrics)




