Clera
Senior Voice AI Engineer
About this role
Description: Founding engineer on a conversational AI team, responsible for the real-time voice layer from speech input to spoken responses. Focus on making natural voice interactions reliable in production with an emphasis on end-to-end latency. Requirements: Minimum 5 years of production software experience, with at least 2 years in voice, speech, or real-time audio systems. Proven experience in building end-to-end real-time voice pipelines, including speech recognition and synthesis. Proficient in Python or TypeScript, with hands-on experience in audio stacks like LiveKit or Twilio Media Streams. Experience debugging audio at the frame level and building LLM evaluation harnesses. Strong written English skills for asynchronous communication; experience with SIP or LLM orchestration is a plus. Benefits: Annual compensation of $96,000 USD, regardless of location. Fully remote position with core team overlap from 13:00 to 17:00 UTC.