Digital Zone
Senior Site Reliability Engineer (Performance and Scalability)
About this role
Description: Build the platform's scalability foundation to handle campaign-level traffic spikes. Establish load and failure testing as a standard practice across engineering teams. Own SLOs, error budgets, and observability stack for various services. Harden Postgres and AWS infrastructure for performance and availability. Lead incident response and drive systemic fixes into design and planning. Partner with engineering teams to enhance their self-sufficiency in scaling. Requirements: 5+ years in SRE, platform, or backend engineering with experience in large-scale systems. Proven track record of scaling systems during traffic spikes and implementing testing programs. Deep experience with AWS and strong understanding of Postgres performance. Proficiency in observability tooling, infrastructure-as-code, and scripting languages like Go or TypeScript. Strong communication skills and a calm approach to incident management. Benefits: Immediate impact on a high-growth business. Competitive compensation packages. Opportunity to work with top regional talent from leading companies.