LOG IN
SIGN UP
Canary Wharfian - Online Investment Banking & Finance Community.
Sign In
or continue with e-mail and password
Forgot password?
Don't have an account?
Join Canary Wharfian
or continue with e-mail and password
By signing up, you agree to our Terms & Conditions and Privacy Policy.

Data Engineer - Assistant Vice President

ExperiencedNo visa sponsorship
Citi logo

at Citi

Bulge Bracket Investment Banks

Posted 4 days ago

No clicks

**Data Engineer - Assistant Vice President in Pune, India**. Design, develop, and optimize real-time data pipelines using Python, PySpark, Spark SQL, and Databricks (must-haves). Lead cloud integrations (AWS or GCP), ensure data quality, and govern pipelines. Mentor junior engineers and collaborate cross-functionally. Must have 3+ years in Databricks, strong SQL, and distributed computing skills. Proven expertise in data warehousing concepts and cloud platforms. Exceptional problem-solving, communication, and adaptability skills required. Full-time hybrid role.

Compensation
Not specified

Currency: Not specified

City
Not specified
Country
India

Full Job Description

Data Engineer - Assistant Vice President

Apply (opens in new window)
Save

Job Req Id:

26963387

Location(s):

Pune, Maharashtra, India

Job Type:

Hybrid

Posted:

Jul. 31, 2026

Discover your future at Citi

Working at Citi is far more than just a job. A career with us means joining a team of approximately 219,000 dedicated people from around the globe. At Citi, youll have the opportunity to grow your career, give back to your community and make a real impact.

Job Overview

The Job description is given below - 

A C12 Data Engineer is expected to:

  • Act as a Subject Matter Expert (SME): Provide technical guidance on pipeline architecture, distributed computing, and cloud-native integrations.
  • Drive Best Practices: Champion code quality, comprehensive automated testing, CI/CD automation, and rigorous data governance standards.

Key Responsibilities

  • Data Pipeline Architecture & Development: Design, develop, and deploy scalable batch and real-time end-to-end data pipelines (ETL/ELT) using Python, PySpark, Spark SQL, and Databricks(Must have).
  • Cloud Infrastructure Integration: Deploy and maintain Databricks workspaces on cloud environments (AWS or GCP). Manage secure integrations with cloud storage (S3/GCS), access controls (IAM), secrets management (Vault/KMS), and serverless query engines.
  • Performance Optimization & Tuning: Diagnose and resolve performance bottlenecks in Spark clusters, SQL queries, and Databricks jobs. Optimize storage layouts using Delta Lake properties (e.g., Z-Ordering, partitioning, and vacuuming).
  • Data Quality & Governance: Implement automated data validation frameworks, data quality monitoring, and metadata management solutions utilizing Databricks Unity Catalog to ensure strict compliance with internal data governance policies and external financial regulations (such as BCBS 239).
  • Technical Leadership & Mentorship: Act as a technical lead within an Agile/Scrum environment. Lead peer code reviews, enforce coding standards, and mentor junior and mid-level data engineers (C10/C11).
  • DevOps & CI/CD: Establish and maintain automated CI/CD pipelines (using Jenkins, GitLab, or GitHub Actions) for packaging and deploying data engineering artifacts (dbt, Spark jobs, Databricks workflows).
  • Collaboration: Partner with Data Science teams to operationalize machine learning models, and work with business intelligence developers to build efficient semantic layers for reporting.

Technical Qualifications (Must-Haves)

  • Programming Languages: Strong, production-grade proficiency in Python (including standard libraries, pandas, and testing frameworks like pytest) and advanced SQL (including window functions, CTEs, and query optimization).
  • Distributed Computing: Deep hands-on experience with Apache Spark (PySpark) for processing multi-terabyte datasets in a distributed cluster environment.
  • Unified Lakehouse Platforms: Minimum 3 years of hands-on experience developing within Databricks. Expert knowledge of Delta Lake ACID transactions, Delta Live Tables (DLT), Unity Catalog, and Databricks Workflows is required.
  • Cloud Platforms: Extensive experience deploying Databricks within either Amazon Web Services (AWS) or Google Cloud Platform (GCP). Proficiency in cloud-native components (AWS S3, EC2, IAM, EMR, Athena, Redshift OR GCP GCS, Compute Engine, IAM, Dataproc, BigQuery) is a strict requirement.
  • Data Modeling: Solid understanding of data warehousing concepts, including dimensional modeling (Star and Snowflake schemas), slow-changing dimensions (SCDs), and Medallion (Bronze/Silver/Gold) architecture design.

Preferred Qualifications & Certifications

  • Databricks Certifications: Databricks Certified Data Engineer Professional or Databricks Certified Associate Developer for Apache Spark.
  • Cloud Certifications: AWS Certified Solutions Architect / AWS Certified Data Engineer, or Google Cloud Professional Data Engineer.

Professional Competencies & Soft Skills

  • Problem-Solving: Exceptional analytical and troubleshooting skills to resolve complex performance and data consistency issues in distributed systems.
  • Communication: Excellent verbal and written communication skills. Ability to articulate complex technical architectures to non-technical business stakeholders.
  • Collaborative Mindset: Proactive team player who thrives in a diverse, global, cross-functional engineering environment.
  • Adaptability: Ability to prioritize work, pivot quickly in response to changing business requirements, and master new technologies as they emerge.

Thanks !

------------------------------------------------------

Job Family Group:

Technology

------------------------------------------------------

Job Family:

Applications Development

------------------------------------------------------

Time Type:

Full time

------------------------------------------------------

Most Relevant Skills

Please see the requirements listed above.

------------------------------------------------------

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

------------------------------------------------------

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi( opens in new window).

View Citis EEO Policy Statement( opens in new window) and the Know Your Rights( opens in new window) poster.

Apply (opens in new window)
Save

Data Engineer - Assistant Vice President

Compensation

Not specified

City: Not specified

Country: India

Citi logo
Bulge Bracket Investment Banks

4 days ago

No clicks

at Citi

ExperiencedNo visa sponsorship

**Data Engineer - Assistant Vice President in Pune, India**. Design, develop, and optimize real-time data pipelines using Python, PySpark, Spark SQL, and Databricks (must-haves). Lead cloud integrations (AWS or GCP), ensure data quality, and govern pipelines. Mentor junior engineers and collaborate cross-functionally. Must have 3+ years in Databricks, strong SQL, and distributed computing skills. Proven expertise in data warehousing concepts and cloud platforms. Exceptional problem-solving, communication, and adaptability skills required. Full-time hybrid role.

Full Job Description

Data Engineer - Assistant Vice President

Apply (opens in new window)
Save

Job Req Id:

26963387

Location(s):

Pune, Maharashtra, India

Job Type:

Hybrid

Posted:

Jul. 31, 2026

Discover your future at Citi

Working at Citi is far more than just a job. A career with us means joining a team of approximately 219,000 dedicated people from around the globe. At Citi, youll have the opportunity to grow your career, give back to your community and make a real impact.

Job Overview

The Job description is given below - 

A C12 Data Engineer is expected to:

  • Act as a Subject Matter Expert (SME): Provide technical guidance on pipeline architecture, distributed computing, and cloud-native integrations.
  • Drive Best Practices: Champion code quality, comprehensive automated testing, CI/CD automation, and rigorous data governance standards.

Key Responsibilities

  • Data Pipeline Architecture & Development: Design, develop, and deploy scalable batch and real-time end-to-end data pipelines (ETL/ELT) using Python, PySpark, Spark SQL, and Databricks(Must have).
  • Cloud Infrastructure Integration: Deploy and maintain Databricks workspaces on cloud environments (AWS or GCP). Manage secure integrations with cloud storage (S3/GCS), access controls (IAM), secrets management (Vault/KMS), and serverless query engines.
  • Performance Optimization & Tuning: Diagnose and resolve performance bottlenecks in Spark clusters, SQL queries, and Databricks jobs. Optimize storage layouts using Delta Lake properties (e.g., Z-Ordering, partitioning, and vacuuming).
  • Data Quality & Governance: Implement automated data validation frameworks, data quality monitoring, and metadata management solutions utilizing Databricks Unity Catalog to ensure strict compliance with internal data governance policies and external financial regulations (such as BCBS 239).
  • Technical Leadership & Mentorship: Act as a technical lead within an Agile/Scrum environment. Lead peer code reviews, enforce coding standards, and mentor junior and mid-level data engineers (C10/C11).
  • DevOps & CI/CD: Establish and maintain automated CI/CD pipelines (using Jenkins, GitLab, or GitHub Actions) for packaging and deploying data engineering artifacts (dbt, Spark jobs, Databricks workflows).
  • Collaboration: Partner with Data Science teams to operationalize machine learning models, and work with business intelligence developers to build efficient semantic layers for reporting.

Technical Qualifications (Must-Haves)

  • Programming Languages: Strong, production-grade proficiency in Python (including standard libraries, pandas, and testing frameworks like pytest) and advanced SQL (including window functions, CTEs, and query optimization).
  • Distributed Computing: Deep hands-on experience with Apache Spark (PySpark) for processing multi-terabyte datasets in a distributed cluster environment.
  • Unified Lakehouse Platforms: Minimum 3 years of hands-on experience developing within Databricks. Expert knowledge of Delta Lake ACID transactions, Delta Live Tables (DLT), Unity Catalog, and Databricks Workflows is required.
  • Cloud Platforms: Extensive experience deploying Databricks within either Amazon Web Services (AWS) or Google Cloud Platform (GCP). Proficiency in cloud-native components (AWS S3, EC2, IAM, EMR, Athena, Redshift OR GCP GCS, Compute Engine, IAM, Dataproc, BigQuery) is a strict requirement.
  • Data Modeling: Solid understanding of data warehousing concepts, including dimensional modeling (Star and Snowflake schemas), slow-changing dimensions (SCDs), and Medallion (Bronze/Silver/Gold) architecture design.

Preferred Qualifications & Certifications

  • Databricks Certifications: Databricks Certified Data Engineer Professional or Databricks Certified Associate Developer for Apache Spark.
  • Cloud Certifications: AWS Certified Solutions Architect / AWS Certified Data Engineer, or Google Cloud Professional Data Engineer.

Professional Competencies & Soft Skills

  • Problem-Solving: Exceptional analytical and troubleshooting skills to resolve complex performance and data consistency issues in distributed systems.
  • Communication: Excellent verbal and written communication skills. Ability to articulate complex technical architectures to non-technical business stakeholders.
  • Collaborative Mindset: Proactive team player who thrives in a diverse, global, cross-functional engineering environment.
  • Adaptability: Ability to prioritize work, pivot quickly in response to changing business requirements, and master new technologies as they emerge.

Thanks !

------------------------------------------------------

Job Family Group:

Technology

------------------------------------------------------

Job Family:

Applications Development

------------------------------------------------------

Time Type:

Full time

------------------------------------------------------

Most Relevant Skills

Please see the requirements listed above.

------------------------------------------------------

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

------------------------------------------------------

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi( opens in new window).

View Citis EEO Policy Statement( opens in new window) and the Know Your Rights( opens in new window) poster.

Apply (opens in new window)
Save