Relocantly← All jobs

🇩🇪 Berlin, Germany · 6h ago

Machine Learning Engineer

distil labs

LinkedInEnglish-friendly
Apply on LinkedIn →Get jobs like this daily
Level: Senior, 5+ yearsWhat are we building?distil labs is a platform that fine-tunes task-specific small language models automatically. Customers use it to swap the general-purpose LLM in their agentic system for a smaller, purpose-built one: same quality on the task, at 50-80% lower cost and with lower latency. Under the hood, we take the production traces collected by the customer, generate synthetic training data from them, train a small model that matches frontier-model quality on the narrow task, and deploy it to an OpenAI-compatible endpoint in the cloud, on-prem, or at the edge.About The RoleWe are looking for a Machine Learning Engineer to own the lifecycle of every model we ship: synthetic data generation, distributed training and evaluation, and deployment to production endpoints. A large part of the job is working with our science team to develop new methods in knowledge distillation, synthetic data generation and model self-improvement, and turning the ones that prove out into production stages that make every customer’s model better. As part of a small team, you will have an outsized impact on technical decisions, product direction, and engineering culture.What you’ll doOwn the model lifecycle end to end: from the customer’s raw traces, through synthetic data generation and validation, distributed fine-tuning and evaluation, to a deployed endpoint.Work with our science team to develop new methods in knowledge distillation, synthetic data generation and model self-improvement, then turn the ones that work into production pipeline stages that run for every customer.Build the evaluation and benchmarking harness that tells us whether a change actually made our models better, and run the experiments that answer that question.Run and optimise distributed fine-tuning workloads (HuggingFace, PyTorch, DDP/FSDP, LoRA) across cloud and on-prem GPU clusters.Operate the compute fabric it all runs on (Argo Workflows, Kubernetes and similar) and the secure, multi-tenant serving layer (vLLM, FastAPI) that holds low-latency SLAs under load.What you’ll bring5+ years building and shipping machine learning systems in production, with hands-on ownership of training and evaluation, and deploying models.Comfort reading a machine learning paper and turning it into a working pipeline, and comfort working with scientists on the method itself rather than only on its implementation.Experimental rigour: you can design an evaluation that tells you whether a change improved the model, and you trust the measurement over the intuition.Deep proficiency in Python for machine learning. Our stack is built on the HuggingFace ecosystem with PyTorch.Proven track record running distributed training jobs (DDP/FSDP, DeepSpeed, Ray or Kubeflow) and a working knowledge of where they break.Hands-on experience operating Kubernetes, Argo Workflows or similar orchestration systems in production, alongside containerisation and cloud services (AWS, GCP or Azure).Bonus points forHands-on experience with knowledge distillation, synthetic data generation or model compression.Judgement about training data: what makes a training set good, and how to spot a synthetic example that will hurt more than it helps.Familiarity with inference optimisation techniques (quantisation, sparsity, compilation) and serving engines (Triton, TensorRT-LLM, vLLM).Depth in infrastructure-as-code (Terraform, Pulumi), observability stacks (Prometheus, Grafana, Datadog) and GPU cost optimisation.Experience running hybrid or on-prem GPU clusters and high-speed storage (NVMe, Infiniband).Experience contributing to academic research, ideally with publications in machine learning or related fields.Contributions to open-source ML infrastructure projects.Why should you join?Real ownership: own the lifecycle end to end and choose the best tools for the job.Science that ships: work directly with our research group and see new methods reach production customers in weeks, not years.Mission with impact: make advanced AI accessible to teams that lack massive GPUs or ML expertise. Our models already run in production for defence, cybersecurity, edtech and robotics customers.Early-stage upside: competitive salary, VSOP/ESOP and a real say in where the company goes.Remote-first, Europe-centric: work from anywhere in EU time zones, with regular Berlin offsites.Why now?Enterprises are moving from general-purpose cloud LLMs to smaller, faster, private models, and the methods and infrastructure to train and serve those models securely are still being invented. Join us at the ground floor and shape the systems that will power the next generation of AI products.Who are we?We are a small team full of deep technical expertise: engineers and researchers who built ML systems and infrastructure at places like Amazon, Delivery Hero, Five AI and ING. Meet the whole team on our homepage.distil labs is an equal-opportunity employer. We value diversity and do not discriminate on the basis of race, religion, colour, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Sourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.