🇩🇪 Berlin, Germany · 2h ago
ML Data Engineer
XpertDirect
LinkedInseniorEnglish-friendly
ML Data EngineerBerlin, Germany — HybridAI SaaS | ML Data Engineering | Data Infrastructure | Machine Learning | Production AIOur client, a growing AI SaaS company based in Berlin, is looking for an ML Data Engineer to build the pipelines, datasets, and data infrastructure used to train, evaluate, and operate machine-learning systems in production.You'll work at the intersection of Data Engineering and Machine Learning, ensuring ML teams have reliable, reproducible, and production-ready data throughout the model lifecycle.What You'll Work On• Build scalable data pipelines using Python and SQL• Develop distributed data-processing workflows using Apache Spark• Orchestrate training and data workflows using Apache Airflow• Build reliable datasets for model training and evaluation• Develop ingestion and transformation pipelines for structured and unstructured data• Implement automated data-quality checks and validation• Track datasets, experiments, and model artefacts using MLflow• Build and operate data workloads on AWS• Improve reproducibility across ML training and evaluation pipelines• Monitor data freshness, completeness, consistency, and quality• Troubleshoot production data issues affecting ML systems• Optimise pipelines for performance and growing data volumes• Collaborate with ML Engineers and Data Scientists to move models from experimentation into productionCore Skills• 3+ years in Data Engineering, ML Data Engineering, ML Infrastructure, or a similar role• Python• SQL• Apache Spark• Apache Airflow• AWS• MLflow• Data quality and validation• Experience building production data pipelinesNice to HavePyTorch / TensorFlowDatabricksKafkaSnowflakedbtFeature engineeringFeature storesData versioningGreat Expectations / SodaDocker / KubernetesTerraformModel training pipelinesData lineage and observabilitySourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.