Keep IT Simple
Keep IT Simple
Keep IT Simple (KIS) is a Silicon Valley-based IT services and solutions provider serving organizations across California and beyond. Established in 1988, the company helps clients address complex technology needs through cybersecurity, virtualization, cloud solutions, and network infrastructure consulting. Its work spans industries including logistics, healthcare, and cybersecurity, with an emphasis on practical IT guidance, cost-conscious solutions, and responsive customer service.

Senior Databricks Platform Engineer

Build and operate secure Databricks platforms for a global insurance and asset management company. The role covers cloud infrastructure, governance, reliability, and AI workloads.

Description

  • Design and operate enterprise Databricks platforms across cloud environments.
  • Deploy Databricks workspaces, clusters, and platform components on Azure and/or AWS.
  • Maintain scalable, secure, and highly available Databricks environments.
  • Set platform standards, architecture patterns, and operational practices.
  • Manage Unity Catalog, Delta Lake, and workspace governance.
  • Handle platform lifecycle management, upgrades, and capacity planning.
  • Provision infrastructure with Terraform and automate deployments and configuration management.
  • Design networking, private connectivity, and secure integrations.
  • Implement backup, disaster recovery, and high availability.
  • Improve cloud resource use and manage platform costs.
  • Apply enterprise security controls and compliance requirements.
  • Configure role-based access and least-privilege models.
  • Manage secrets, key vault integrations, and encryption standards.
  • Set up monitoring, audit logs, and governance controls.
  • Support regulatory and security audits.
  • Monitor platform health, availability, and performance.
  • Build observability for infrastructure and data workloads.
  • Troubleshoot platform, networking, and workload problems.
  • Create operational runbooks and support processes.
  • Lead root-cause analysis and remediation.
  • Work with data, AI, and cloud infrastructure teams.
  • Advise on platform architecture and operational excellence.
  • Create engineering standards and reusable automation patterns.
  • Mentor engineers and promote cloud engineering practices.
  • Support Databricks AI and machine learning environments.
  • Enable MLflow experimentation, model tracking, and deployment.
  • Support Mosaic AI, Vector Search, Feature Store, and RAG architectures.
  • Assist teams building generative AI and large language model solutions.
  • Develop infrastructure and governance for AI workloads.
  • Assess new Databricks AI capabilities and recommend adoption strategies.

Requirements

  • Bachelor’s degree in computer science, information technology, engineering, or a related field.
  • At least five years of cloud, platform, or infrastructure engineering experience.
  • At least three years of hands-on experience deploying and administering Databricks environments.
  • Experience supporting enterprise cloud platforms on Azure and/or AWS.
  • Experience designing secure, scalable, highly available cloud solutions.
  • Skills in Databricks administration, Unity Catalog, Delta Lake, workspace management, cluster configuration and optimization, workflows, and job scheduling.
  • Knowledge of Microsoft Azure and Amazon Web Services (AWS).
  • Terraform experience.
  • Experience with Azure DevOps or GitHub Actions.
  • CI/CD pipeline experience.
  • Scripting skills in PowerShell, Python, or Bash.
  • Knowledge of RBAC, private endpoints, VNET/VPC design, encryption and key management, identity federation, and secrets management.
  • Experience with Databricks monitoring, Azure Monitor, CloudWatch, Log Analytics, performance tuning, and capacity management.
  • Preferred: Databricks Certified Platform Administrator, Databricks Certified Data Engineer Professional, Azure Administrator Associate, Azure Solutions Architect, or AWS Solutions Architect certification.
  • Preferred experience implementing enterprise AI and machine learning platforms.
  • Preferred experience supporting regulated industries such as insurance or financial services.
  • Desired familiarity with Mosaic AI, MLflow, Databricks Model Serving, Vector Search, Feature Store, RAG, large language models, AI governance and responsible AI practices, and AI-enabled data engineering patterns.

Benefits

  • Contractor engagement model (PJ).
  • Hybrid work with three days per week in person at the Pinheiros, São Paulo office.

Related Jobs

Aleph Alpha

Senior AI Researcher, Foundation Model Pre-Training — Aleph Alpha, Heidelberg

Aleph Alpha
51 – 200 Employees
Artificial IntelligenceB2B

Lead architecture and large-scale pre-training work for Aleph Alpha’s European foundation models, shaping training methods across thousands of GPUs. Build and refine PyTorch training recipes with the Heidelberg-based hybrid team.

Open