Seneca Holdings
Seneca Holdings
501 – 1,000 Employees
ConsultingHealthcareLogistics
Seneca Holdings is the investment arm of the Seneca Nation, developing and managing profitable businesses that strengthen long-term economic self-sufficiency for the Nation. Its diversified portfolio includes federal government contracting, environmental solutions, and health-related services, with a focus on sustainable opportunities that support the Seneca community and future generations.

Senior Data Engineer - Databricks and AWS (Remote)

Build scalable Databricks and AWS data pipelines for federal government contracting. Deliver analytics-ready data while meeting FISMA High and multi-tenant compliance standards.

Description

  • Build and maintain scalable data pipelines in an AWS-hosted Databricks E2 environment
  • Source data from relational databases, APIs, external providers, and real-time streams
  • Create pipelines that cleanse, transform, and aggregate data for analytics and reporting
  • Develop and manage Bronze, Silver, and Gold layers in a Databricks medallion architecture
  • Prepare source-to-target mappings and perform unit testing
  • Write advanced SQL aggregations and investigate data anomalies, quality problems, and inconsistencies
  • Implement real-time and near-real-time ingestion with AWS DMS and other AWS-native services
  • Develop processing solutions with Python or R, Spark, PySpark, and Pandas
  • Manage code, version control, and deployments through GitLab and CI/CD pipelines
  • Work with cross-functional teams in an Agile delivery environment
  • Maintain compliance with FISMA High and multi-tenant security, governance, and regulatory requirements
  • Use AI automation tools to support pipeline development, testing, validation, and delivery

Requirements

  • Demonstrated hands-on experience as a Data Engineer or Databricks Developer building production data pipelines
  • Strong knowledge of Databricks on AWS, including E2 architecture, cluster configuration, job orchestration, and workspace administration
  • Practical experience using Apache Spark, PySpark, and Pandas for distributed data processing at scale
  • Strong Python programming ability; R experience is advantageous
  • Experience using Databricks Auto Loader for scalable incremental file ingestion
  • Hands-on experience with AWS Database Migration Service for change data capture and real-time replication
  • Strong command of Delta Lake and Delta tables, including schema evolution, time travel, OPTIMIZE, Z-ORDER, and VACUUM
  • Experience working with Amazon RDS and other relational sources in extraction and integration workflows
  • Advanced SQL expertise covering complex joins, window functions, aggregations, and query performance tuning
  • Understanding of medallion architecture and modern data lakehouse principles
  • Experience using GitLab and CI/CD pipelines for automated testing, builds, and deployment of data engineering code
  • Experience contributing to Agile or Scrum projects

Benefits

  • Competitive compensation
  • Medical, dental, vision, life, and disability coverage
  • Optional critical illness, hospital, and accident coverage
  • Health savings and flexible spending accounts
  • 401(k) retirement plan
  • Paid leave programs
  • Flexible work-life balance
  • Professional development opportunities
  • Performance and recognition programs
  • Collaborative workplace
  • Financial and non-financial benefits for Seneca Nation members

Related Jobs