Gcore
Python Inference Engineer
About this role
Description: Build and improve the inference layer of the Gcore Inference platform. Integrate and operate inference frameworks such as vLLM, SGLang, NVIDIA Dynamo, and TensorRT-LLM. Bring new language and multimodal models into production. Improve inference latency, throughput, memory use, GPU utilization, and cost efficiency. Debug performance and reliability issues across model code, inference frameworks, GPU execution, networking, and Kubernetes. Collaborate with various teams to turn inference improvements into reliable product features. Contribute to open-source inference projects when appropriate. Requirements: 5+ years of experience writing reliable, well-tested production code. Strong Python skills and experience designing production systems. Hands-on experience with PyTorch and deploying machine learning models. Experience with Linux, Docker, and Kubernetes. Experience in at least one relevant area: distributed systems, GPU computing, ML runtimes, model optimization, or cluster scheduling. Ability to debug complex problems across software, infrastructure, and hardware. Strong sense of developer experience and genuine interest in inference engineering. Good communication and collaboration skills. Benefits: Competitive compensation. Flexible working hours and hybrid or remote options. Work from anywhere in the world for up to 45 days per year. Private medical insurance for you and your family.* Extra paid vacation and sick leave days.* Support for life’s important moments and celebrations. Language courses to help you connect and grow. Modern, welcoming offices with snacks, drinks, and entertainment.* Team sports and social activities.* *Benefits may vary depending on your location.