
at J.P. Morgan
Bulge Bracket Investment BanksPosted 10 days ago
No clicks
**Senior Data Scientist** - Design, develop, and maintain scalable data pipelines with Airflow for trusted reporting and analysis. - Build and manage data products (dimensions, models, marts) with clear ownership and documentation. - Deliver metrics and datasets for AI adoption measurement using SQL and dbt; manage definition changes. - Implement data quality controls and governance, with experience in Databricks and Spark/PySpark. - Collaborate with cross-functional teams to define requirements and interpret metrics across PDLC/SDLC. - Mentor peers and improve engineering best practices; no formal people management. - Bring 3+ years of production data solution building expertise, strong data modeling skills, and delivery ownership. - Experience with metrics evolution, SQL, dbt, Python/PySpark, and orchestration pipelines (Airflow).
- Compensation
- Not specified
- City
- Not specified
- Country
- United States
Currency: Not specified
Full Job Description
Location: OH, United States
Job Description
We are seeking a Data Science Senior Associate focused on building and operating resilient datasets, pipelines, and reusable metrics that support hypothesis-driven analyses and experiments across the product development lifecycle (PDLC)
In this role, you will be hands-on in designing, developing, and maintaining data products that are reliable, observable, and well-documentedenabling partners across product, engineering, and analytics to measure whats driving value, where friction exists, and how operating-model changes impact outcomes as teams adopt more agentic ways of working. Youll contribute to engineering standards and help raise the quality bar through strong delivery and collaboration.
Job Responsibilities
- Build and operate scalable batch/streaming pipelines with SLAs, monitoring, and incident response participation (as needed).
- Create and maintain trusted data products (dimensions, event models, marts) with clear ownership and documentation.
- Deliver metrics and feature-ready datasets for AI adoption/productivity measurement; manage definition changes over time.
- Implement data quality and governance controls (validation, reconciliation, lineage, access, retention, auditability).
- Orchestrate workflows in Airflow (or equivalent), including backfills and retries.
- Model/transform data using SQL and dbt (or equivalent) for trusted reporting and repeatable measurement.
- Write production-grade Python/PySpark with testing, performance tuning, and maintainable design.
- Partner with cross-functional stakeholders to define requirements, success criteria, and metric interpretation across finance, PDLC/SDLC, and AI tool logs.
- Contribute to engineering best practices (version control, code review, CI/CD, runbooks) and improve observability and cost/performance.
- Mentor peers through reviews, documentation, and knowledge sharing (no formal people management).
Required Qualifications
- Bachelors degree in Computer Science, Engineering, or equivalent practical experience.
- 3+ years building production data solutions; strong ownership and delivery.Strong engineering fundamentals (OOP, testing, development lifecycle).Strong data modeling skills (dimensional, normalized, event-based).Experience with Databricks and/or Spark/PySpark.Strong SQL; experience with dbt (or equivalent) and building testable data codebases.Experience operating orchestration pipelines (Airflow or equivalent).Proven ability to build and maintain reliable metrics as sources/definitions evolve.Effective delivery in ambiguous, multi-stakeholder environments.
Preferred Qualifications
- Experience with modern lakehouse/warehouse patterns and broader cloud data platforms (e.g., Databricks, Snowflake).
- Experience with BI/semantic layers and metrics management practices.
- Exposure to experimentation or hypothesis-driven analytics approaches (e.g., measurement design to support tests, rollouts, and pre/post evaluation); deep causal specialization not required.
- Experience improving observability (data freshness/SLA monitoring, lineage, alerting) and contributing to operational maturity (runbooks, incident follow-ups).
Data Scientist, Senior Associate
Compensation
Not specified
City: Not specified
Country: United States
ExperiencedNo visa sponsorship**Senior Data Scientist** - Design, develop, and maintain scalable data pipelines with Airflow for trusted reporting and analysis. - Build and manage data products (dimensions, models, marts) with clear ownership and documentation. - Deliver metrics and datasets for AI adoption measurement using SQL and dbt; manage definition changes. - Implement data quality controls and governance, with experience in Databricks and Spark/PySpark. - Collaborate with cross-functional teams to define requirements and interpret metrics across PDLC/SDLC. - Mentor peers and improve engineering best practices; no formal people management. - Bring 3+ years of production data solution building expertise, strong data modeling skills, and delivery ownership. - Experience with metrics evolution, SQL, dbt, Python/PySpark, and orchestration pipelines (Airflow).
Full Job Description
Location: OH, United States
Job Description
We are seeking a Data Science Senior Associate focused on building and operating resilient datasets, pipelines, and reusable metrics that support hypothesis-driven analyses and experiments across the product development lifecycle (PDLC)
In this role, you will be hands-on in designing, developing, and maintaining data products that are reliable, observable, and well-documentedenabling partners across product, engineering, and analytics to measure whats driving value, where friction exists, and how operating-model changes impact outcomes as teams adopt more agentic ways of working. Youll contribute to engineering standards and help raise the quality bar through strong delivery and collaboration.
Job Responsibilities
- Build and operate scalable batch/streaming pipelines with SLAs, monitoring, and incident response participation (as needed).
- Create and maintain trusted data products (dimensions, event models, marts) with clear ownership and documentation.
- Deliver metrics and feature-ready datasets for AI adoption/productivity measurement; manage definition changes over time.
- Implement data quality and governance controls (validation, reconciliation, lineage, access, retention, auditability).
- Orchestrate workflows in Airflow (or equivalent), including backfills and retries.
- Model/transform data using SQL and dbt (or equivalent) for trusted reporting and repeatable measurement.
- Write production-grade Python/PySpark with testing, performance tuning, and maintainable design.
- Partner with cross-functional stakeholders to define requirements, success criteria, and metric interpretation across finance, PDLC/SDLC, and AI tool logs.
- Contribute to engineering best practices (version control, code review, CI/CD, runbooks) and improve observability and cost/performance.
- Mentor peers through reviews, documentation, and knowledge sharing (no formal people management).
Required Qualifications
- Bachelors degree in Computer Science, Engineering, or equivalent practical experience.
- 3+ years building production data solutions; strong ownership and delivery.Strong engineering fundamentals (OOP, testing, development lifecycle).Strong data modeling skills (dimensional, normalized, event-based).Experience with Databricks and/or Spark/PySpark.Strong SQL; experience with dbt (or equivalent) and building testable data codebases.Experience operating orchestration pipelines (Airflow or equivalent).Proven ability to build and maintain reliable metrics as sources/definitions evolve.Effective delivery in ambiguous, multi-stakeholder environments.
Preferred Qualifications
- Experience with modern lakehouse/warehouse patterns and broader cloud data platforms (e.g., Databricks, Snowflake).
- Experience with BI/semantic layers and metrics management practices.
- Exposure to experimentation or hypothesis-driven analytics approaches (e.g., measurement design to support tests, rollouts, and pre/post evaluation); deep causal specialization not required.
- Experience improving observability (data freshness/SLA monitoring, lineage, alerting) and contributing to operational maturity (runbooks, incident follow-ups).




