StackYak
Senior AI Inference Engineer
About this role
Description: StackYak is an early-stage company building AI infrastructure software that integrates compute, GPU infrastructure, networking, and inference. The role involves determining how models run on available hardware, focusing on precision, quantization, parallelism, and serving runtime. Requirements: Proven experience running large language models in production and serving across multiple GPUs and nodes. Strong background in measuring and defending performance metrics like latency, throughput, and utilization. Proficiency in Python for production-level tooling and debugging Linux systems. Experience with NVIDIA and/or AMD inference environments. Benefits: Competitive compensation with meaningful equity. Opportunity to work in a small, senior team with direct access to founders and a collaborative environment.