🇫🇷 France · 10h ago
Founding Machine Learning Engineer
Hub
LinkedInleadEnglish-friendly
Turn raw multi-camera footage into the 3D, verified datasets that frontier robotics labs train on.About HubHub is one of the fastest-growing data companies, providing real-world training data to the largest AI and robotics companies. Based in San Francisco and backed by Y Combinator and top VCs, our mission is to advance embodied AGI through real-world data and research.The roleYou own Hub's ML pipeline end to end, from raw multi-camera capture in the field to the dataset a frontier robotics lab trains on, side by side with our core ML team. You are the technical counterpart for each lab program, and your pipeline decides what we are allowed to ship.What you'll own- 3D and stereo vision on capture rigs we build ourselves: multi-camera RGB, RGB-D and IMU, calibration, stereo and metric depth, SLAM and trajectories in a world frame, 3D reconstruction of the scene and the manipulation.- Hand tracking across synchronised cameras and fine-grained manipulation, plus annotation at scale against demanding customer taxonomies. Every human verdict becomes a training label.- Quality control as an ML problem: vision-language models as judges, thresholds per customer, human reviewers where models fall short. You own the eval sets and the call on what we trust.- Customer pipelines on AWS: extend our shared modules, build what a new spec demands, and keep throughput and cost per processed hour under control.- Egocentric video with narration across languages: speech recognition, translation, chapter segmentation, QA against each customer's taxonomy.- The technical relationship with each lab program: specs, delivery format, first samples and feedback loops.- What comes next: tactile and teleoperation data today, and robot learning on our own data in the mid term.You might be a fit if- 3D and robotics stack: camera calibration, stereo depth, SLAM or visual-inertial odometry, multi-view geometry, and multimodal sensor data (video, depth, IMU, time sync).- Perception models in production: detection, tracking, hand and body pose, depth. Strong PyTorch: you read architectures, compare and benchmark them, and train or fine-tune when needed. Robotics perception counts most.- 3+ years of applied ML in production, ideally from a top engineering school. A PhD counts toward the years and is a strong plus. Less experience is fine for outliers: the bar is what you've built.- Infra, MLOps and scaling: AWS (S3, GPU instances, Kubernetes), multi-GPU training, model serving, data pipelines at scale, experiment tracking and labeling tools.- You've owned a delivery under hard deadlines, and you explain things clearly to researchers, field teams and reviewers.- Agentic engineering as a craft: a custom harness, and agents that verify their own work through tests, training runs and evals.- Security awareness: you know how a Linux or cloud pipeline gets attacked (DDoS, intrusions, leaked keys, exposed services) and how to harden it.Nice to have- VLMs as judges or annotators in production.- Egocentric vision, IMUs, MCAP, ROS.- World models, video generation, VLAs or robot foundation models.- An applied PhD or published work in 3D vision or robot learning.Stack- PyTorch, CUDA, multi-GPU training and distributed inference. Throughput matters as much as accuracy.- Multi-camera geometry: intrinsics, camera-to-IMU extrinsics, fisheye rectification, stereo depth, SLAM and trajectories in a world frame.- Multi-stream HEVC video plus high-rate IMU per recording, hardware-synchronised and calibrated, delivered as MCAP.- Frontier VLMs and speech models behind one interface, plus open-weight judges we serve with vLLM and fine-tune on our own data.- AWS: S3 for bytes, GPU instances and Kubernetes for compute, Postgres for state. Cost per processed hour is an engineering target.- Every threshold traces to a customer requirement. Every quarantine carries a code, evidence and an owner.Why Hub- Small core team by design, San Francisco pace. Urgent things get handled when they come up.- One owner per project, with 1 to 3 numbers that say whether it's working.- Everyone runs AI agents, not only engineers. Uncapped frontier models, shared skills and knowledge base.- Constant, proactive communication: Slack, huddles, short stand-ups. No black box.- Team across San Francisco, São Paulo and Paris, where our research lab is opening.What we offer- $90,000 to $120,000 yearly salary.- Stock options between 0.05% and 0.2%.- Your own GPU budget.- Based in Paris, in our office opening soon. Hybrid: ideally most days on site, at least one day a week (or one week a month if you live outside Paris).- Direct work with the founders and with the biggest AI labs.How we hire1. A first call with our team: your motivation, culture fit and the basics.2. A take-home assignment over a weekend (6 to 8 hours is enough) on our public Hugging Face datasets (huggingface.co/Hubdata).- Run a hand detection and tracking system on our egocentric stereo RGB-D set- Report what works and what fails with numbers- Fix one real issue, and use the depth- Pose or IMU data for something beyond 2D detection.- Comparing the same system on our GoPro and iPhone sets is a plus.- Send your code with a short write-up or slides.3. A 60-minute debrief with one of our ML engineers on your assignment: how you reasoned about the data and the sensors, what you measured and what you would do next.4. A final conversation with a founder.5. An answer within 48 hours.How to applySend a short application in your own words:- Your CV- The achievement you're most proud of- Your GitHub and Hugging Face links- The story behind itSourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.