Clera
Robotics Data Infrastructure Engineer
About this role
About the Role This role sits at the intersection of robotics and data infrastructure at a well-funded, early-stage robotics company. You'll be responsible for building reliable pipelines and storage systems that make large volumes of robot telemetry and sensor data usable for engineering and ML teams. Your work will directly enable faster iteration and safer robotic systems — from raw sensor ingestion all the way through to training-ready datasets and real-time analytics. You'll join a cross-functional team of robotics engineers, software engineers, and data scientists in a fast-paced, on-site environment in Los Angeles, CA. This is a high-impact, hands-on role with broad scope at a company building at the frontier of physical AI and robotics. Please note: Visa sponsorship is not available for this role.
What You'll Do
- Design and build scalable data pipelines to ingest and process robot telemetry and sensor data (camera, LiDAR, IMU, and more).
- Implement storage solutions and schemas that support analytics, model training, and data replay.
- Ensure data quality, validation, and lineage across ingestion and transformation stages.
- Optimize latency and throughput for both real-time and batch processing use cases.
- Instrument observability, monitoring, and alerting for data flows and infrastructure.
- Collaborate closely with robotics engineers and data scientists to translate platform needs into production-grade implementations.
- Productionize ETL/ELT workflows with CI/CD and automated testing.
- Troubleshoot and resolve production incidents affecting data availability or correctness. What We're Looking For Required:
- 3+ years of hands-on experience building data infrastructure or engineering pipelines specifically for robotics sensor data — this is a dealbreaker requirement.
- Proven experience designing, building, and maintaining data ingestion, processing, and storage pipelines for sensor data (e.g., camera, LiDAR, IMU).
- Strong fundamentals in distributed systems, databases, and data pipeline design.
- Proficiency in Python and/or C++ for building data tooling and pipelines.
- Hands-on experience with cloud data platforms and distributed processing tools — e.g., AWS or GCP, Kafka or Pub/Sub, Spark or Flink, Airflow.