SSC HR Solutions
Senior Data Engineer (Spark)
About this role
Description: Builds and runs large scale batch data pipelines, ensuring efficiency and cost-effectiveness as data volumes increase. Utilizes Apache Spark for managing large scale data processing jobs, including performance tuning for reliable execution. Supports data modeling for analytics and organizes data in a lakehouse for downstream usability. Handles automated scheduling, monitoring, and data quality checks, collaborating with platform and product teams on end-to-end data flows. Requirements: Strong hands-on experience with Apache Spark, including performance tuning (not just through managed notebooks). Experience in data modeling for analytics and organizing data in a lakehouse. Proficient in automated scheduling, monitoring, and data quality checks. Ability to work with platform and product teams on end-to-end data flows. Familiarity with Apache Iceberg or other open table formats. Experience with Trino or similar query engines. Experience with on-premises or self-managed clusters. Benefits: Opportunity to work with cutting-edge technologies in data engineering. Collaborative work environment with cross-functional teams.