Site Reliability Engineer
Maintain and improve the reliability, observability, and operational maturity of enterprise platforms through systematic engineering discipline.
Quix is looking for a Site Reliability Engineer who approaches operational problems with an engineering mindset — building systems, tooling, and processes that prevent incidents rather than only responding to them. This role is for someone with deep experience in observability, incident management, capacity planning, and reliability engineering in distributed, production environments.
What You’ll Do
- Define and track service-level indicators and service-level objectives for critical platform services.
- Build and maintain observability infrastructure: metrics, distributed tracing, structured logging, and alerting.
- Lead post-incident analysis and drive systematic improvements that reduce mean time to recovery.
- Establish runbooks, escalation paths, and operational playbooks for production platform teams.
- Conduct capacity planning, load testing, and failure mode analysis to proactively identify reliability risks.
- Collaborate with engineering teams on production readiness reviews, deployment gates, and reliability requirements.
- Identify and drive down toil through automation, infrastructure improvements, and process redesign.
What We’re Looking For
- Substantial production experience in an SRE, platform engineering, or infrastructure reliability role.
- Deep expertise with observability tooling: Prometheus, Grafana, Datadog, OpenTelemetry, or equivalents.
- Experience managing distributed systems reliability: load balancing, failover, circuit breakers, and retry logic.
- Strong incident management experience: on-call discipline, root cause analysis, and post-mortem culture.
- Proficiency with scripting and automation for operational tasks in Python, Bash, or Go.
- Ability to communicate reliability requirements and tradeoffs clearly to engineering and product teams.
Nice to Have
- Experience with chaos engineering practices and failure injection tooling.
- Background in SLO and SLA management in enterprise or regulated service environments.
- Experience with multi-region or multi-cloud reliability architectures.
Platform reliability is not an afterthought — it is a core commitment to clients who depend on Quix systems for operational continuity. The SRE function protects that commitment by building systems that fail gracefully, recover quickly, and improve systematically, reducing operational risk across every deployment.
Submit for Site Reliability Engineer
The role is pre-selected. Resume and mobile phone are required so the intake can support verification when connected.