Databento
Databento
Databento provides market data infrastructure through APIs for real-time and historical financial data. Its platform delivers normalized information across futures, options, equities, and other asset classes, including live streams, historical datasets, corporate action details, and full exchange order book replay. Data is sourced from colocation sites to support low-latency access, while developer-focused tools and integrations for Python and C++ help teams build and operate market data applications. Databento serves more than 3,000 firms and startups with flexible access and pricing options.

Site Reliability Engineer — Remote, United States

Maintain reliability, performance, and observability for Databento’s financial market-data platform. Improve deployments, incident response, and backend infrastructure across API and platform services.

Description

  • Own uptime, SLAs, and SLOs for API and platform services
  • Establish reliability and operational standards for developers
  • Build and maintain logging, metrics, and tracing observability
  • Design and operate highly available deployments and containerized environments
  • Profile and optimize Python applications for throughput, latency, and cost
  • Diagnose production issues at the operating-system level with strace, perf, eBPF, ss, and gdb
  • Strengthen deployment processes and CI/CD workflows
  • Join the on-call rotation, lead incident response, and facilitate post-incident reviews
  • Determine necessary fixes and carry projects from initial concept through completion
  • Contribute to petabyte-scale data processing, customer management and billing, and query systems that power APIs

Requirements

  • Mid-level or senior individual contributor
  • Professional experience in SRE, DevOps, or backend engineering, ideally within a trading firm, technology company, or high-growth startup
  • Practical experience with observability tools for logs, metrics, and traces, including Prometheus, OpenTelemetry, VictoriaMetrics, Jaeger, Logstash, Loki, or Vector
  • Experience with containerization and highly available deployment using technologies such as Docker, Podman, Docker Compose, Docker Swarm, Kubernetes, or k3s
  • Advanced Python skills, including application development and performance tuning
  • Proficiency with Linux debugging and profiling tools such as strace, perf, eBPF, ss, and gdb
  • Demonstrated measurable results in a recent position
  • Experience applying alerting and incident-response practices is advantageous
  • Familiarity with configuration management or infrastructure-as-code tools such as Ansible or Terraform is beneficial
  • Experience with HTTP benchmarking, load testing, and capacity planning is a bonus
  • Database schema design and query optimization experience is desirable
  • Strong communication skills and a dependable work ethic in a remote environment
  • Interest in financial data or algorithmic trading

Benefits

  • Equal employment opportunity and protection against discrimination
  • Workplace accommodations available upon request
  • Option to opt out of AI-powered Talent Matching

Related Jobs

Knowtion Health

Remote Talent Acquisition Manager

Knowtion Health

Oversee recruiting systems, requisitions, analytics, and contingent workforce operations at Knowtion Health. Help support scalable hiring processes for a growing healthcare company.

Open