OpenRouter
OpenRouter
OpenRouter is an artificial intelligence platform that gives developers and teams a unified way to access and work with large language models from different providers. Its model catalog includes options such as Mistral Small 3 and DeepSeek R1 Distill Qwen 32B, supporting applications that prioritize speed, accuracy, and flexible model selection. OpenRouter also offers tools for coding and gaming, with potential applications across areas including marketing, finance, and education.

Site Reliability Engineer, AI Provider Operations — OpenRouter (Remote, US)

Operate reliability systems for OpenRouter’s AI providers, including observability, failover, and incident response. Automate endpoint quality checks, capacity management, and launch readiness.

Description

  • Develop monitoring for provider endpoints and track latency, throughput, errors, uptime, and output correctness.
  • Define service-level objectives (SLOs) for provider tiers and configure actionable alerts.
  • Improve degraded-endpoint detection and coordinate automatic traffic failover with the routing team.
  • Manage on-call provider incidents from triage and mitigation through partner communication, postmortems, and follow-up actions.
  • Produce provider scorecards and SLO reports.
  • Provide technical escalation support for provider endpoint issues.
  • Develop continuous canaries and evaluations to catch otherwise hidden quality regressions.
  • Automate provider operations, including endpoint disablement, capacity changes, deprecations, and rate-limit adjustments.
  • Create endpoint load-testing tools to assess launch readiness and initial traffic.
  • Report to the Provider Operations Manager.

Requirements

  • At least four years in SRE, production engineering, or infrastructure roles supporting high-traffic, customer-facing systems.
  • Strong observability skills, including metrics, tracing, logs, SLOs, error budgets, and actionable alerting.
  • Software engineering ability and a preference for building tools over following runbooks.
  • Experience with TypeScript and/or Python.
  • Knowledge of distributed-systems failure modes, including timeouts, retries, backpressure, and partial outages.
  • Ability to lead incidents calmly and communicate clearly with external partners under pressure.
  • Understanding of LLM inference concepts, or willingness to learn about streaming, tool calling, prompt caching, throughput and latency tradeoffs, and provider API differences.
  • Experience at an inference provider, model lab, GPU cloud, or API gateway or CDN company is a plus.
  • Experience with TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, or Vercel is a plus.
  • Background in routing, load balancing, or traffic management systems is a plus.
  • Experience with evaluations or synthetic monitoring for machine-learning systems is a plus.

Related Jobs