CrewAI
Software Engineer, Infrastructure & Reliability
About this role
Description: Build and operate the platform infrastructure for CrewAI's cloud and enterprise deployments across AWS, Azure, and GCP. Focus on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Requirements: Strong infrastructure/platform engineering experience in production SaaS environments. Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services. Experience with ECS and/or Kubernetes; Helm experience is a strong plus. Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production. Strong debugging instincts across app, infra, network, deploy, and dependency layers. Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access. Ability to write reliable automation in Python, Ruby, Go, Bash, or similar. Calm, rigorous approach to incidents, rollbacks, migrations, and production change management. Benefits: Opportunity to work with cutting-edge AI technology and multi-agent systems. Contribute to a platform that powers 300M+ agent executions per month across thousands of companies.