Salesforce
Salesforce
Salesforce develops cloud-based software centered on a customer relationship management platform that brings company data and teams together through an integrated, AI-driven system. Its products support sales, customer service, and marketing workflows, with solutions designed for businesses ranging from small and medium-sized enterprises to larger organizations. Salesforce also provides industry-focused tools, Trailhead learning resources, and AppExchange apps that extend the capabilities of its CRM platform.

Senior Lead Site Reliability Software Engineer

Lead reliability engineering for Salesforce’s AI-powered CRM platform, scaling secure infrastructure and automation for critical customer workflows.

Description

  • Own reliability roadmaps for major product areas and evolve architectures into highly available, globally scalable systems
  • Collaborate on infrastructure strategy, system design, capacity planning, bottleneck analysis, and security
  • Develop and maintain infrastructure-as-code and deployment tooling
  • Help scale and deploy AI and machine learning infrastructure
  • Create systems and processes for product testing and release management
  • Build automation for customer and partner deployments, upgrades, diagnostics, and operations
  • Strengthen identity and access controls, networking, workload isolation, secret management, deployment security, and policy enforcement
  • Apply AI tools to automate operational work and accelerate infrastructure delivery
  • Identify and prioritize architectural issues across distributed backend systems
  • Implement application performance, infrastructure, and data-model improvements
  • Set engineering standards, guardrails, and team practices that reduce regressions
  • Work with customer, product, and engineering teams to understand needs and make practical tradeoffs
  • Mentor engineers and share expertise in backend architecture and system design

Requirements

  • Experience operating mission-critical production systems and scaling rapidly growing products
  • Advanced proficiency with Kubernetes, Terraform or OpenTofu, and AWS, Google Cloud, or Azure
  • Ability to write and review production-grade Golang, TypeScript, or Python
  • Deep knowledge of distributed systems and debugging interactions among microservices, databases, and AI agents
  • Experience working on a senior team that includes Principal engineers
  • Proven ability to use modern AI development tools
  • At least five years of experience in SRE, production engineering, or backend engineering focused heavily on operations and infrastructure
  • Advanced proficiency with Golang, GraphQL, and PostgreSQL
  • Experience with RDS, Redis or ElastiCache, and EKS
  • Expertise in backend engineering, including data modeling, database performance, state management, API design, transactionality, concurrency, memory management, fault tolerance, and scaling
  • Repeatedly owning foundational systems work in complex distributed environments
  • Strong written, verbal, and technical design communication skills
  • An advanced computer science degree or equivalent practical experience is advantageous
  • Experience in regulated or highly secured industries, especially public-sector environments with compliance and data-sovereignty requirements, is advantageous
  • Familiarity with classified or limited-connectivity environments, including Department of Defense Impact Level 6, is advantageous
  • Advanced knowledge of microservice orchestration and durability patterns, including Temporal and service mesh technologies, is advantageous
  • Deep understanding of networking, security, and identity management across major cloud providers is advantageous
  • Experience building AI products for supply chain, logistics, manufacturing, or operational workflows is advantageous
  • Bachelor’s degree in Computer Science required; master’s degree preferred
  • Exposure to Temporal, Istio, or Typesense
  • Experience with CI/CD platforms, particularly Jenkins and Spinnaker
  • Exposure to supply chain, logistics, or manufacturing industries
  • Familiarity with the Salesforce platform
  • Experience with workflow engines

Benefits

  • Paid time-off programs
  • Medical coverage
  • Dental coverage
  • Vision coverage
  • Mental health resources
  • Paid parental leave
  • Life insurance
  • Disability insurance
  • 401(k) retirement plan
  • Employee stock purchase program
  • Some roles may qualify for incentive compensation
  • Some roles may qualify for equity awards
  • Reasonable accommodations throughout the application and recruiting process

Related Jobs

Avalon Healthcare Solutions

Senior AI Engineer, Florida Remote | Avalon Healthcare Solutions

Avalon Healthcare Solutions

Build and lead production AI systems—including agentic workflows, retrieval-augmented generation, and automation—for Avalon Healthcare Solutions’ diagnostic intelligence platform. Help improve healthcare operations and clinical outcomes through reliable, governed AI solutions.

Open
Bitdeer Group

AI Engineer, New Graduate

Bitdeer Group
201 – 500 Employees
ConstructionConsultingLogistics

Build and scale AI cloud infrastructure, GPU and Kubernetes services, model-serving systems, and AI Agent platforms. Contribute across Bitdeer AI’s full-stack infrastructure for global AI workloads.

Open