Luma
Software Engineer, Inference
About this role
Description: Own the integration of new model architectures into the inference engine. Collaborate with research, engineering, and infrastructure teams to optimize model efficiency and deployments. Build internal tools to measure, profile, and track inference jobs and workflows. Automate, test, and maintain inference services for maximum uptime and reliability. Manage and optimize inference workloads across clusters and hardware providers. Build scheduling systems to optimize GPU resource usage while meeting SLOs. Requirements: Strong Python and system-architecture skills. Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar. Experience with queues, scheduling, traffic control, and fleet management at scale. Proficiency in Linux, Docker, and Kubernetes, including orchestration and deployment. Familiarity with Redis and S3-compatible storage. Benefits: Opportunity to work on large-scale inference systems and cutting-edge technologies. Engage in a collaborative environment across multiple teams to drive innovation.