Runpod
Senior ML Systems Engineer, Inference
About this role
Description: Runpod is seeking a Senior ML Systems Engineer, Inference to enhance LLM inference performance on their AI Developer Cloud platform. The role involves measuring, diagnosing, and improving inference performance, impacting customer experience directly. Requirements: 5+ years of professional system engineering experience. Deep experience with vLLM, SGLang, or similar serving engines at scale. Strong software engineering skills in Python, with experience in performance-critical codebases. Understanding of LLM inference performance drivers: batching, memory, parallelism, and trade-offs. Familiarity with inference optimization techniques like quantization and distributed serving. Rigor in benchmarking and performance analysis, with GPU profiling tool experience. Ability to communicate results clearly in writing. Benefits: Competitive base pay ranging from $150,000 to $220,000, adjusted based on experience and location. Meaningful equity in a fast-growing company with stock options for all team members. Generous medical, dental, and vision plans. Flexible PTO for work-life balance. Remote work-first environment with collaborative team culture. $1,200 Home Office & Equipment Stipend to set up an ideal workspace.