Capgemini
Capgemini
Capgemini is a consulting and technology services company that helps organizations transform and manage their operations. Its work spans industries including aerospace, automotive, banking, healthcare, and logistics, with services covering cloud, cybersecurity, data and artificial intelligence, and enterprise management. Capgemini also focuses on innovation and sustainability as part of its approach to digital transformation. The company hires across a broad range of professions and experience levels, with an emphasis on building an innovative and diverse workforce.

Senior Data Engineer (Remote, Canada) at Capgemini

Build scalable data-processing applications and ETL pipelines that support Capgemini’s digital transformation services. Optimize distributed workloads for analytics, reporting, and machine learning.

Description

  • Create, enhance, and support scalable data-processing applications with Apache technologies and SQL
  • Develop high-throughput ETL and data-transformation pipelines in Python and Scala
  • Build and tune batch and near-real-time data-processing workflows
  • Handle large datasets across Parquet, Hive, and distributed storage platforms
  • Create performant SQL queries and improve execution plans for scale and efficiency
  • Define data models, partitioning approaches, and storage architectures
  • Partner with data engineers, software engineers, data scientists, and product teams
  • Establish automated testing, monitoring, and deployment workflows for data applications
  • Maintain data quality, reliability, security, and governance requirements

Requirements

  • Bachelor’s or master’s degree in computer science, engineering, or a related technical discipline
  • At least five years of professional experience in software engineering or data engineering
  • Practical expertise with Apache technologies, SQL, Python, and Scala
  • Background developing applications for distributed data processing
  • Solid knowledge of Hadoop ecosystem tools such as Hive, Parquet, and HDFS
  • Experience creating and tuning ETL pipelines for large-scale data
  • Strong SQL development and query-tuning capabilities
  • Working knowledge of Linux systems and shell scripting
  • Experience with source-control tools such as Git
  • Preferred experience includes AWS, distributed computing, large-scale data architecture, Airflow orchestration, and data lake or lakehouse environments

Benefits

  • 12 to 25 vacation days based on grade
  • Paid company holidays
  • Personal days
  • Sick leave
  • Medical, dental, and vision coverage, or coordination with provincial healthcare in Canada
  • Retirement savings programs, such as RRSPs in Canada
  • Life and disability insurance
  • Employee assistance programs
  • Additional benefits determined by local policy and eligibility
  • Potential variable compensation, bonuses, or commissions may be available

Related Jobs