OpenRouter
Site Reliability Engineer, Provider Operations
About this role
Description: OpenRouter is an AI routing and infrastructure layer for enterprises to manage and optimize large language models. The Site Reliability Engineer (SRE) will ensure operational health of provider supply, managing requests and endpoints effectively. Requirements: 4+ years in SRE, production engineering, or infrastructure roles with high-traffic systems. Proficient in observability tools and practices, including metrics, tracing, logs, and alerting. Strong software engineering skills in TypeScript and/or Python. Experience with distributed systems and incident management. Willingness to learn about LLM inference and provider APIs. Benefits: Opportunity to work at the forefront of AI technology and infrastructure. Be part of a team that handles high-volume requests and ensures system reliability.