Replit
Staff Site Reliability Engineer
About this role
Description: Join Replit's Site Reliability Engineering (SRE) team to ensure the reliability, scalability, and performance of infrastructure serving millions of developers. Implement automation and establish best practices to maintain high availability and improve system reliability. Requirements: 8-10 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering). Strong programming skills in Python or Go, with a focus on high-quality, well-tested code. Deep understanding of distributed systems and experience with Kubernetes and cloud-native technologies. Proven track record in designing and maintaining monitoring and observability solutions. Strong incident management skills and experience leading complex incident responses. Familiarity with infrastructure as code tools (Terraform, Pulumi) and configuration management. Excellent communication skills and ability to mentor engineers at various levels. Benefits: Competitive Salary & Equity. 401(k) Program with a 4% match (US Only). Health, Dental, Vision, and Life Insurance. Short Term and Long Term Disability. Paid Parental, Medical, Caregiver Leave. Flexible Time Off (FTO) + Holidays. Commuter Benefits (In-Office & US Only). Monthly Wellness Stipend. Autonomous Work Environment. In Office Set-Up Reimbursement (In-Office Only). Quarterly Team Gatherings. In Office Amenities (In-Office Only).