Pattern Bioscience
Pattern Bioscience
Pattern Bioscience is a biotechnology and healthcare company developing rapid, culture-free diagnostics for drug-resistant bacterial infections. Its technology combines single-cell analysis of microorganisms with machine learning to produce clinically actionable results within hours, helping healthcare professionals identify appropriate treatments sooner. The company’s work addresses antimicrobial resistance by supporting faster, more informed decisions that may improve patient outcomes and reduce avoidable healthcare costs.

Senior Data Engineer, Data and Analytics – Pune Hybrid

Lead Pattern Bioscience’s advertising optimization data systems, including the feature store and machine learning inference workflows. Maintain reliable, quality-controlled bid and budget updates for marketplace advertising APIs.

Description

  • Build, deploy, and support scalable automated batch pipelines that ingest diverse sources into the lakehouse
  • Manage and enhance Airflow orchestration for daily and 15-minute multi-DAG workflows, covering dependencies, retries, backfills, full refreshes, and safe reruns
  • Develop and optimize analytical SQL using wide joins, window functions, incremental merges, warehouse sizing, and query profiling
  • Add features and labels to the feature store while maintaining established leakage-prevention and data-completeness standards
  • Coordinate SageMaker model training and batch inference through Airflow, including datasets, S3 and Parquet transfers, training images, instance sizing, predictions, and metrics
  • Create data-auditing approaches and enforce blocking quality checks before external writes
  • Diagnose and resolve large-scale data-processing issues, investigate failures, and participate in on-call triage
  • Protect idempotency, new-data detection, action validation, invalidation, and audit trails for marketplace updates
  • Partner with data scientists, advertising strategists, and platform engineering teams
  • Convert ROAS goals, budget pacing, playbook rules, and advertising strategy into dependable data models and pipelines
  • Take ownership of data quality in assigned domains and coordinate issue resolution with data infrastructure teams
  • Mentor data engineers, provide technical guidance, and review SQL and DAG changes

Requirements

  • Bachelor’s degree in data science, data analytics, information management, computer science, information technology, a related discipline, or equivalent professional experience
  • At least four years of overall professional experience
  • At least four years of hands-on SQL and Python development experience, including advanced SQL
  • At least three years of experience building production data pipelines on modern data architectures
  • At least two years of experience with cloud data warehouses such as Snowflake, Amazon Redshift, or BigQuery
  • Professional experience with a workflow orchestrator; Airflow is strongly preferred
  • Experience scheduling machine learning training and batch inference through SageMaker or a comparable platform
  • Working understanding of applied machine learning concepts, including dataset construction, leakage, evaluation metrics, and model lifecycle operations
  • Practical AWS experience, including at least S3 and IAM, plus familiarity with columnar file formats
  • Proven ownership of data quality through testing, monitoring, alerting, and root-cause analysis
  • Strong software engineering and scripting practices, including version control, code review, and modular development
  • Excellent communication and cross-functional collaboration skills
  • Ability to lead and mentor data engineering teams
  • Preferred experience includes digital advertising, advanced Snowflake, time-series data, big data, data-quality frameworks, distributed data platforms, additional AWS services, data governance, and AI coding agents

Benefits

  • Benefits include paid time off, insurance, and competitive compensation
  • Paid time off
  • Insurance

Related Jobs