Coretek
Coretek
Coretek is a Microsoft Azure Expert Managed Service Provider and cloud technology consultancy serving organizations across government, healthcare, manufacturing, and financial services. Its work spans cloud migration, application modernization, DevOps automation, workspace management, cybersecurity, and AI solutions built with Azure OpenAI and machine learning technologies. Coretek also delivers security and observability services using platforms such as Palo Alto and Dynatrace, alongside partnerships involving Imprivata and Citrix. The company’s multidisciplinary team helps clients modernize infrastructure, strengthen security, and develop practical cloud and AI capabilities.

Azure Data Engineer — Hybrid in India

Build and operate dependable Azure batch and streaming pipelines for Coretek. Develop governed data models, quality controls, and production workflows that support business and technical teams.

Description

  • Design, build, and maintain idempotent, observable, and recoverable batch and streaming data pipelines
  • Create analytical data models, including dimensional schemas, semantic layers, and curated marts
  • Connect operational databases, SaaS APIs, files, and event streams while managing schema drift and late-arriving data
  • Embed checks for data freshness, volume, uniqueness, and referential integrity within pipelines
  • Establish alerting and escalation procedures for pipeline failures
  • Manage production pipelines through monitoring, on-call incident response, root-cause analysis, and backfills
  • Improve performance and control costs through partitioning, clustering, file sizing, and warehouse or cluster sizing
  • Use version control, code reviews, CI/CD, automated testing, and infrastructure as code
  • Enforce access controls, PII safeguards, retention, lineage, and audit requirements with security and compliance teams
  • Work with analysts, data scientists, and product engineers to establish durable data contracts
  • Keep data dictionaries, lineage records, and pipeline runbooks current

Requirements

  • At least five years of experience building production data pipelines
  • Strong hands-on Python experience for data engineering, including testing, packaging, and code review
  • Practical command of PySpark DataFrame and SQL APIs, large-scale joins and aggregations, partitioning, shuffle behavior, and Spark UI diagnosis
  • Advanced SQL skills, including window functions, query-plan analysis, and performance tuning
  • Hands-on experience with Azure Data Factory, Databricks, Synapse or Fabric, and ADLS
  • Sound data modeling knowledge covering normalization, star schemas, and slowly changing dimensions
  • Experience with Git-based development workflows and CI/CD delivery
  • Excellent communication skills, including the ability to explain pipeline issues to non-technical stakeholders
  • Outstanding analytical and problem-solving ability, including root-cause analysis
  • Strong experience working with customers in a consultative technical environment
  • Streaming data experience with Kafka or Event Hubs
  • Experience using Delta Lake or Iceberg
  • Experience with Terraform or Bicep for infrastructure as code and Docker or Kubernetes for containerization
  • Experience working in regulated environments involving HIPAA, SOC 2, PCI, or GDPR
  • Experience building machine learning data platforms or supporting feature pipelines
  • Ability to manage multiple client projects and deliver quality work on schedule
  • Experience using Azure DevOps or GitHub for source control and delivery pipelines

Related Jobs

Knowtion Health

Remote Talent Acquisition Manager

Knowtion Health

Oversee recruiting systems, requisitions, analytics, and contingent workforce operations at Knowtion Health. Help support scalable hiring processes for a growing healthcare company.

Open