Cloudlinux
Lead Site Reliability Engineer - Imunify Reliability Platform (remote-only)
About this role
Description: Lead the development of service level indicators (SLIs) for ~70 components of the Imunify360 security suite. Build a telemetry collection system to monitor component performance and alerting mechanisms to detect failures quickly. Define and enforce a structured alert management and escalation process to ensure timely responses to incidents. Requirements: Substantial production-engineering or SRE experience, with a proven track record of defining SLO frameworks. Strong proficiency in Python and familiarity with Go or Rust; experience with time-series telemetry tools like Prometheus and Grafana. Experience in distributed systems debugging, configuration management, and CI at production scale. Excellent written communication skills for async collaboration across teams. Benefits: Opportunities for professional development and learning through challenging projects and mentorship programs. Fully remote work with flexible hours and a healthy work-life balance, including 24 days of paid vacation and unlimited sick leave. Compensation for private medical insurance and reimbursement for co-working and gym/sports expenses. Potential rewards for innovative ideas that can be patented.