Parasail
Senior Site Reliability Engineer
About this role
Description: Join Parasail to build an enterprise-grade inference cloud for AI models, ensuring fast, reliable, and economical service. Work on a global GPU fleet, automating operations and enhancing system reliability and performance. Requirements: Experience in building and operating production infrastructure or distributed systems with a focus on reliability. Strong Linux fundamentals, practical knowledge of networking, storage, and containers, and hands-on experience with Kubernetes in production. Ability to write maintainable software for infrastructure automation and a systematic approach to debugging across various boundaries. Benefits: Opportunity to shape the future of AI infrastructure and make impactful architectural decisions. Work closely with a small, dedicated team on substantial scale problems, directly shipping improvements into production.