Relocantly← All jobs

🇩🇪 Germany · 7h ago

Senior AI Data Engineer

European Tech Recruit

LinkedInseniorEnglish-friendly
Apply on LinkedIn →Get jobs like this daily
Senior AI Data EngineerBerlin, Germany | Hybrid – 3 days onsite weeklyJoin a highly technical AI research and engineering organisation developing the data foundations required to train next-generation machine learning models. As a Senior AI Data Engineer, you will work at the intersection of large-scale data engineering, machine learning and NLP, transforming massive volumes of raw and web-scale data into high-quality training datasets. The Senior AI Data Engineer will take ownership of data processing pipelines covering deduplication, quality scoring, filtering, PII removal, toxicity detection, metadata extraction and evaluation-contamination prevention. The Senior AI Data Engineer will work closely with Research, Engineering and Governance teams to develop new approaches to data quality as model and research requirements evolve, including the use of ML classifiers and LLM-based evaluators. This is an opportunity to work on genuinely large-scale ML data problems where traditional data engineering approaches are not always sufficient.Key ResponsibilitiesDesign, build and scale distributed pipelines that transform web-scale data into high-quality datasets for machine learning model training.Develop data processing workflows covering deduplication, heuristic filtering, PII scrubbing, toxicity removal, metadata extraction and custom data transformations.Build and refine data-quality systems including ML classifiers, LLM-as-a-judge evaluators and automated quality-scoring mechanisms.Develop robust dataset versioning, lineage and provenance processes to ensure training data can be traced and reproduced.Implement leakage and evaluation-contamination detection mechanisms across large-scale data pipelines.Build monitoring, guardrails and alerting systems to identify data-quality regressions before they propagate into downstream training.Conduct dataset coverage analysis and identify gaps in existing corpora, supporting targeted data acquisition strategies.Work with Research and Engineering teams to understand evolving training-data requirements.Collaborate with Legal and Governance teams where required around data privacy, licensing and compliance.Develop internal tools that allow researchers to explore, query and understand large-scale datasets.Required Experience & SkillsDegree in computer science, software engineering, or a related field.At least 5+ years of full-time experience in data processing, machine learning engineering, NLP or a closely related discipline.Proven experience working with massive unstructured text datasets, ideally at billion- or trillion-token scale.Strong hands-on experience with distributed processing frameworks such as Spark, Ray or Flink.Excellent Python skills and experience writing production-grade data-processing software.Practical experience with data quality, data curation, filtering, deduplication or large-scale dataset construction.Experience implementing or working with ML-based data-quality systems, classifiers or LLM evaluation.Understanding of data privacy, PII removal, content filtering and evaluation contamination.Experience with pipeline orchestration technologies such as Airflow, Prefect or Dagster.Strong understanding of dataset versioning, lineage and reproducibility.Experience collaborating across Research, Engineering and/or Governance teams.Desired: Experience with LLM training, fine-tuning or deployment; Familiarity with LLM inference optimisation technologies such as vLLM or SGLang; Experience with Docker and Kubernetes; Experience with web-scale data acquisition or Common Crawl; Experience with data licensing workflows; Contributions to open-source NLP or data-processing projects.Why Apply?Work directly on the data foundations supporting advanced AI and large-scale model training.Solve problems involving trillion-token datasets and web-scale data processing.Combine traditional distributed data engineering with ML, NLP and LLM-based data-quality techniques.Work closely with researchers to define what high-quality training data actually means for evolving AI systems.Opportunity to develop new approaches to data curation, filtering, evaluation and dataset quality rather than simply maintaining established ETL pipelines.Apply now or send a copy of your CV, referencing the title and location, and with a short intro to cw@eu-recruit.com.By applying to this role you understand that we may collect your personal data and store and process it on our systems. For more information please see our Privacy Notice (https://eu-recruit.com/about-us/privacy-notice/)

Sourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.