🇩🇪 Germany · 6h ago

Senior Data Engineer

European Tech Recruit

LinkedInseniorEnglish-friendly
Senior ML Data Processing DeveloperBerlin, Hybrid – 3 days per weekWe’re looking for a Senior ML Data Processing Developer to join a highly technical team working at the intersection of data engineering, machine learning and AI research.You’ll help transform massive, web-scale datasets into high-quality training data used to develop next-generation AI models. This is much more than traditional data engineering—you’ll actively shape data quality through algorithmic filtering, ML-based scoring, data curation and novel processing techniques.As AI research evolves, so will the data challenges. You’ll work on problems where established solutions may not yet exist, helping define new approaches to processing, evaluating, filtering and improving machine-learning datasets at scale.What You'll DoPartner with Research and Engineering teams to design, build, automate and scale large-scale ML data pipelines.Process web-scale text datasets through deduplication, quality scoring, heuristic filtering, toxicity removal, PII scrubbing and metadata extraction.Develop algorithmic approaches to improve the quality, relevance and diversity of training datasets.Build LLM-as-a-Judge evaluators, ML classifiers and human-in-the-loop review workflows.Implement robust data-quality monitoring, guardrails and alerting to identify issues before they reach downstream systems.Establish strong dataset versioning, lineage and provenance while optimising processing for performance and cost.Analyse dataset coverage and identify gaps, then develop targeted strategies for acquiring additional data.Design leakage-detection mechanisms to prevent evaluation and benchmark contamination.Build internal tools that allow researchers to explore, query and understand large datasets efficiently.Work with Legal and Governance teams to ensure data meets privacy, compliance, licensing and governance requirements.Continuously identify opportunities to improve the scale, reliability, quality and efficiency of data processing workflows.What We're Looking ForDegree in Computer Science, Software Engineering or a related technical field.5+ years' experience in data processing, ML engineering, NLP or a closely related discipline.Proven experience handling massive unstructured text datasets, ideally at trillion-token scale.Strong hands-on experience with distributed processing frameworks such as Apache Spark, Ray or Flink.Strong Python skills and experience developing production-grade data-processing systems.Experience with pipeline orchestration tools such as Airflow, Prefect or Dagster.Understanding of PII detection/scrubbing, content safety, toxicity filtering and evaluation-contamination prevention.Ability to work effectively across Research, Engineering, Legal and Governance teams.Strong analytical and problem-solving skills, with the ability to develop solutions to open-ended data challenges.

Sourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.