SecurityScorecard
Senior Site Reliability Engineer
About this role
Description: Senior Site Reliability Engineer responsible for designing and optimizing Kubernetes-based infrastructure and CI/CD systems. Key role in building AI tooling infrastructure and ensuring production reliability through best practices. Requirements: 6+ years in SRE, DevOps, or Infrastructure roles with significant production Kubernetes experience. Hands-on experience with AI/LLM tooling integration and security considerations. Proven success in building CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, etc.). Strong knowledge of Kubernetes internals and managed services (EKS, GKE, AKS). Expertise in Infrastructure as Code (Terraform, Helm, Pulumi) and GitOps. Proficient in Python, Bash, or Go. Familiarity with observability tools (Prometheus, Grafana, Datadog, OpenTelemetry). Production experience with Kafka, Flink, and ClickHouse. Strong communication and cross-team collaboration skills. Benefits: Work in a recognized best workplace with a strong culture of employee engagement. Opportunity to mentor engineers and lead incident response initiatives.