Swoop
Swoop
Swoop develops marketing solutions for pharmaceutical and life sciences organizations, with a focus on direct-to-consumer and healthcare provider engagement. Its work combines AI, data, and privacy-conscious omnichannel campaigns to help connect patients, providers, and brands across channels such as social media and television. The company applies this approach to improving marketing effectiveness while supporting more patient-centered healthcare communication.

Senior Data Engineer at Swoop – Hybrid in California

Lead NimbleRx’s data platform within Swoop’s AI-driven healthcare engagement business. Build compliant pipelines, warehouse models, and performance systems across batch, streaming, and backend services.

Description

  • Own the full data platform, from ingestion and transformation through storage, querying, and access
  • Shape the data platform roadmap as the company’s data requirements expand
  • Develop and improve batch and streaming pipelines with PySpark, EMR, Kinesis, Lambda, and Step Functions
  • Bring data from Postgres, Salesforce, external vendors, and product events into an Iceberg-based lake
  • Establish SCD tables, event models, and shared data standards for engineering and analytics teams
  • Work with product, engineering, analytics, and operations to convert data needs into dependable pipelines
  • Create documentation and tools that support self-service data usage
  • Lead data security and compliance, including audit trails, access controls, and temporary-access processes
  • Improve backend query performance through replica routing, indexing, caching, and I/O instrumentation in Java and Spring services
  • Diagnose and resolve IOPS surges, pipeline outages, schema changes, and delayed data
  • Apply AI to pipeline scaffolding, schema development, and investigations while delivering internal AI tools
  • Coach engineers and analysts in effective use of the data platform

Requirements

  • At least five years of experience delivering production data pipelines and data platforms
  • Advanced Python, PySpark, and SQL skills, including Spark performance tuning at scale
  • Ability and willingness to develop services around the data layer using Java and Spring Boot
  • Practical experience with distributed processing such as Spark and EMR, streaming with Kinesis, and object storage on S3
  • Strong Postgres expertise, including query optimization, indexing, replication, replica routing, and database bottleneck analysis
  • Experience with Iceberg and Trino or comparable technologies
  • Working knowledge of CI/CD practices and Terraform
  • Experience building AI-powered solutions with frontier models or agentic coding tools
  • Demonstrated ability to collaborate across product, operations, and other teams
  • Professional experience treating security and PII/PHI compliance as core engineering responsibilities

Benefits

  • Opportunities for professional development and career growth
  • The company has been recognized as a Best Place to Work; this recognition is not an individual compensation benefit

Related Jobs

Ambush

Senior Machine Learning and AI Engineer

Ambush

Build production-grade generative AI and agent workflows for Ambush’s financial services clients. Develop Python and FastAPI services, SQL data workflows, retrieval-augmented generation, and cloud-based AI systems.

Open