FluidStack
FluidStack
FluidStack provides GPU supercomputing infrastructure for artificial intelligence labs and companies. Its managed cloud platform offers on-demand access to large-scale Nvidia GPU capacity for model training and inference, with deployment and operations support for clusters running technologies such as Kubernetes and Slurm. By managing the underlying hardware and infrastructure, FluidStack helps AI teams scale workloads while focusing on model development, performance, and operational efficiency.

Software Engineer, Ontology (Onsite)

Build ontology, CMDB, DCIM, observability, and automation systems supporting FluidStack’s AI cloud infrastructure. Develop scalable platforms for infrastructure assets, data center operations, and global fleet management.

Description

  • Architect and implement a next-generation CMDB serving as the source of truth for infrastructure assets, network topology, and configuration data
  • Develop DCIM software for rack operations, server and GPU deployment, operating-system installation, quality assurance, and white-screen operations
  • Create asset lifecycle platforms for receiving, racking, inventory, break-fix, and decommissioning processes
  • Engineer monitoring and observability systems that combine BMS, EPMS, and IT-device telemetry with intelligent alerting and incident management
  • Build self-service portals and automation for regional bootstrap, day-two operations, and fleet-scale administration
  • Reduce operational toil by replacing manual processes with workflow automation and self-service tools
  • Implement orchestration systems for incident, problem, and change-management workflows
  • Produce digital-twin visualizations and operational dashboards in partnership with data teams
  • Develop integration layers between internal platforms, vendors, and external systems
  • Work with data center operations, systems, network engineering, security, product, business, support, and operations teams
  • Assess whether platform components should be built internally or purchased
  • Promote CI/CD, infrastructure as code, automated testing, and observability-first engineering
  • Contribute to architecture reviews and technical design decisions
  • Advance engineering quality through code reviews, documentation, and knowledge sharing
  • Design resilient, high-performance systems capable of processing thousands of queries per second
  • Implement monitoring, logging, debugging, and comprehensive error-handling capabilities
  • Plan data migrations and maintain platform dependencies
  • Lead projects from initial concept through deployment and production readiness

Requirements

  • At least three years of professional experience developing production software systems
  • Strong programming ability in Python, Go, or a comparable language
  • Working knowledge of system-design patterns
  • Experience designing REST APIs, data models, and distributed systems
  • Proficiency with relational and NoSQL databases, including PostgreSQL and Redis
  • Practical experience using Docker and infrastructure-as-code tools such as Terraform and Ansible
  • Understanding of CI/CD pipelines and contemporary software development practices
  • Strong knowledge of TCP/IP, DNS, HTTP, and Linux or Unix environments
  • Demonstrated problem-solving skills with attention to scalability, reliability, and operational needs
  • Clear communication skills when working with technical and non-technical stakeholders
  • Experience with CMDB products such as NetBox or Device42, or with asset-management platforms
  • Experience in infrastructure automation, DevOps, or platform engineering
  • Familiarity with workflow orchestration tools such as Temporal, Airflow, or Camunda
  • Knowledge of monitoring and observability technologies including Prometheus, Grafana, or OpenTelemetry
  • Experience working with time-series databases and data visualization
  • Understanding of ITSM frameworks such as ITIL and service-management practices
  • Experience with data center operations, facilities management, or physical infrastructure
  • Contributions to open-source infrastructure projects
  • Bachelor’s degree in Computer Science or equivalent practical experience
  • U.S. work authorization requirements are covered during the application process

Benefits

  • Equity is available for all full-time positions and may include restricted stock units
  • Retirement or pension benefits are provided according to local norms
  • Health, dental, and vision insurance
  • Generous paid time off aligned with local norms
  • Competitive total compensation package
  • Commitment to pay equity and transparency

Related Jobs

Hansa Biopharma

Senior Accountant at Hansa Biopharma - New York City Hybrid

Hansa Biopharma

Senior Accountant responsible for payroll, financial close, reconciliations, internal controls, and public-company reporting at Hansa Biopharma. The role also supports audits, compliance, complex accounting, and finance operations in a growing biotechnology business.

Open