🇪🇸 Spain · 3w ago

Data-Engineer (Mid)

SoccerSolver

LinkedInmidEnglish-friendly
Role MissionOwn the platform's data layer: ensure football data (in-house + Wyscout and future sources) arrives clean, reconciled, updated daily, and available to the ML algorithms and the application. Also reduces reliance on a single person for data.ResponsibilitiesDesign, build, and maintain ingestion and transformation pipelines (in-house + third-party).Solve entity reconciliation (entity resolution / record linkage) across sources with different identifiers. First major project: Wyscout integration.Set up and maintain orchestration of daily processes and incremental/merge load logic.Implement data quality controls, monitoring, and pipeline alerting.Model and optimize the database schema for ML and product use cases.Document pipelines, schemas, and decisions so the team can operate without depending on a single person.Must-have requirementsReal experience building and operating data pipelines in production (not just notebooks/analysis).Advanced SQL and solid relational data modeling (PostgreSQL ideally).Python for data processing.Experience with process orchestration: Airflow / Cloud Composer, or equivalents like Cloud Run jobs + Scheduler, Dagster, Prefect…Experience with entity reconciliation/deduplication or record linkage: fuzzy matching, canonical keys, ambiguity resolution.Data quality mindset: validation, idempotency, failure handling.Nice-to-havesGCP experience: Cloud SQL, BigQuery, Cloud Run, Composer. Not a dealbreaker: a strong AWS/Azure profile adapts in 1-2 weeks.Experience working alongside ML teams (serving features, feature stores).Good infra/data practices: CI, data tests, schema versioning.

Sourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.