LOG IN
SIGN UP
Canary Wharfian - Online Investment Banking & Finance Community.
Sign In
or continue with e-mail and password
Forgot password?
Don't have an account?
Join Canary Wharfian
or continue with e-mail and password
By signing up, you agree to our Terms & Conditions and Privacy Policy.

Lead Software Engineer - Databricks, ML, AWS

ExperiencedNo visa sponsorship
J.P. Morgan logo

at J.P. Morgan

Bulge Bracket Investment Banks

Posted 6 days ago

No clicks

**Lead Software Engineer - Databricks, ML, AWS, 5+ years** Architect and deliver high-throughput, low-latency data pipelines using Databricks & Apache Spark. Establish lakehouse patterns with Delta Lake, ensuring performance at scale. Drive team AI adoption for improved code quality and delivery speed. Manage Databricks clusters and orchestrate jobs with Databricks Workflows & AWS eventing. Build reusable libraries and implement CI/CD for data projects. Requires 8+ years' experience in software/data engineering, Databricks expertise, and a strong security mindset. Lead and mentor team, driving code quality and technical best practices.

Compensation
Not specified USD

Currency: $ (USD)

City
Not specified
Country
United States

Full Job Description

Location: Plano, TX, United States

We have an exciting and rewarding opportunity for you to take your software engineering career to the next level. 

As a Lead Software Engineer, Machine Learning and Cloud at JPMorgan Chase within the Corporate Technology- Consumer & Community Bank Finance group, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor and lead, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firms business objectives.

Job Responsibilities:

  • Lead architecture and delivery of high-throughput, low-latency data pipelines using Databricks and Apache Spark (Core, SQL, Structured Streaming).
  • Establish lakehouse patterns with Delta Lake (ACID transactions, schema evolution, time travel, Z-ordering, compaction) and ensure performance at scale.
  • Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
  • Own Databricks cluster strategy and setup: runtime selection, autoscaling, driver/executor sizing, Spark configs, unit scripts, cluster policies, pools, and instance profiles.
  • Orchestrate jobs with Databricks Workflows; integrate with AWS eventing and orchestration as needed.
  • Design secure data ingestion and transformation frameworks leveraging Databricks services: Design delta or unmanaged tables, Create tasks for data, ingestion process, Create DAGs using Airflow to orchestrate creation of trusted and refined data.
  • Enforce data quality, lineage, and governance using Unity Catalog and/or Glue Catalog; embed expectations and validation into pipelines.
  • Drive Spark performance engineering: partitioning strategies, file sizing, AQE, broadcast joins, shuffle tuning, caching, spill/memory control, and job right-sizing to optimize cost.
  • Build reusable libraries, frameworks, and APIs in Python and/or Java; oversee unit, integration, and data validation testing.
  • Implement CI/CD for data projects (Git-based workflows), Terraform Infrastructure deployments environment promotion, and automated deployments; champion engineering standards and code reviews. 

 

Required qualifications, capabilities, and skills:

  • Formal training or certification on software engineering concepts and 5+ years applied experience. 
  • 8+ years of professional software/data engineering experience, including substantial production work with Spark on Databricks or EMR.
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
  • Strong proficiency in Python and/or Java for data processing, platform tooling, and automation.
  • Hands-on Databricks expertise (Delta Lake, Unity Catalog, Workflows, Repos/notebooks, SQL Warehouses).
  • Proven track record architecting and operating ETL/ELT pipelines (batch and streaming), with schema design/evolution, SLAs, and reliability engineering.
  • Deep skills in Spark performance tuning and Databricks cluster setup/optimization.
  • Strong SQL and analytics data modeling (dimensional/star schema; lakehouse best practices).
  • CI/CD and automation tooling for data (Git workflows, artifact management) and testing frameworks (pytest, JUnit).
  • Security-first mindset: roles/instance profiles, secret management, encryption-at-rest/in-transit, and network controls. 

 

Preferred qualifications, capabilities, and skills:

  • Experience with Delta Live Tables and advanced governance (catalogs, grants, auditing) in Databricks.
  • AWS networking knowledge (VPC, subnets, routing, security groups) and data egress controls.
  • Experience with Terraform for Infra deployments
  • Cost optimization experience: autoscaling strategies, spot vs on-demand, auto-termination, storage layouts and compaction.
  • Observability for data systems (freshness/completeness metrics, lineage, SLAs, alerting).
  • Drive databricks performance tuning through liquid clustering or partitioning keys, familiarity with Airflow, Genie, Streamlit and React
  • Demonstrated leadership in code quality, reviews, testing strategy, CI/CD, and technical mentorship; excellent communication with stakeholders.
Carry out critical tech solutions across multiple technical areas as an integral part of an agile team.

Lead Software Engineer - Databricks, ML, AWS

Compensation

Not specified USD

City: Not specified

Country: United States

J.P. Morgan logo
Bulge Bracket Investment Banks

6 days ago

No clicks

at J.P. Morgan

ExperiencedNo visa sponsorship

**Lead Software Engineer - Databricks, ML, AWS, 5+ years** Architect and deliver high-throughput, low-latency data pipelines using Databricks & Apache Spark. Establish lakehouse patterns with Delta Lake, ensuring performance at scale. Drive team AI adoption for improved code quality and delivery speed. Manage Databricks clusters and orchestrate jobs with Databricks Workflows & AWS eventing. Build reusable libraries and implement CI/CD for data projects. Requires 8+ years' experience in software/data engineering, Databricks expertise, and a strong security mindset. Lead and mentor team, driving code quality and technical best practices.

Full Job Description

Location: Plano, TX, United States

We have an exciting and rewarding opportunity for you to take your software engineering career to the next level. 

As a Lead Software Engineer, Machine Learning and Cloud at JPMorgan Chase within the Corporate Technology- Consumer & Community Bank Finance group, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor and lead, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firms business objectives.

Job Responsibilities:

  • Lead architecture and delivery of high-throughput, low-latency data pipelines using Databricks and Apache Spark (Core, SQL, Structured Streaming).
  • Establish lakehouse patterns with Delta Lake (ACID transactions, schema evolution, time travel, Z-ordering, compaction) and ensure performance at scale.
  • Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
  • Own Databricks cluster strategy and setup: runtime selection, autoscaling, driver/executor sizing, Spark configs, unit scripts, cluster policies, pools, and instance profiles.
  • Orchestrate jobs with Databricks Workflows; integrate with AWS eventing and orchestration as needed.
  • Design secure data ingestion and transformation frameworks leveraging Databricks services: Design delta or unmanaged tables, Create tasks for data, ingestion process, Create DAGs using Airflow to orchestrate creation of trusted and refined data.
  • Enforce data quality, lineage, and governance using Unity Catalog and/or Glue Catalog; embed expectations and validation into pipelines.
  • Drive Spark performance engineering: partitioning strategies, file sizing, AQE, broadcast joins, shuffle tuning, caching, spill/memory control, and job right-sizing to optimize cost.
  • Build reusable libraries, frameworks, and APIs in Python and/or Java; oversee unit, integration, and data validation testing.
  • Implement CI/CD for data projects (Git-based workflows), Terraform Infrastructure deployments environment promotion, and automated deployments; champion engineering standards and code reviews. 

 

Required qualifications, capabilities, and skills:

  • Formal training or certification on software engineering concepts and 5+ years applied experience. 
  • 8+ years of professional software/data engineering experience, including substantial production work with Spark on Databricks or EMR.
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
  • Strong proficiency in Python and/or Java for data processing, platform tooling, and automation.
  • Hands-on Databricks expertise (Delta Lake, Unity Catalog, Workflows, Repos/notebooks, SQL Warehouses).
  • Proven track record architecting and operating ETL/ELT pipelines (batch and streaming), with schema design/evolution, SLAs, and reliability engineering.
  • Deep skills in Spark performance tuning and Databricks cluster setup/optimization.
  • Strong SQL and analytics data modeling (dimensional/star schema; lakehouse best practices).
  • CI/CD and automation tooling for data (Git workflows, artifact management) and testing frameworks (pytest, JUnit).
  • Security-first mindset: roles/instance profiles, secret management, encryption-at-rest/in-transit, and network controls. 

 

Preferred qualifications, capabilities, and skills:

  • Experience with Delta Live Tables and advanced governance (catalogs, grants, auditing) in Databricks.
  • AWS networking knowledge (VPC, subnets, routing, security groups) and data egress controls.
  • Experience with Terraform for Infra deployments
  • Cost optimization experience: autoscaling strategies, spot vs on-demand, auto-termination, storage layouts and compaction.
  • Observability for data systems (freshness/completeness metrics, lineage, SLAs, alerting).
  • Drive databricks performance tuning through liquid clustering or partitioning keys, familiarity with Airflow, Genie, Streamlit and React
  • Demonstrated leadership in code quality, reviews, testing strategy, CI/CD, and technical mentorship; excellent communication with stakeholders.
Carry out critical tech solutions across multiple technical areas as an integral part of an agile team.