🇫🇷 Paris, France · 1d ago
Lead LLM Engineer
Leonar
LinkedInleadEnglish-friendly
Licorne Society a été missionné par une startup IA en pleine croissance pour les aider à trouver leur Lead LLM Engineer.What You Will OwnYou will be responsible for one thing:Make our AI outputs reliable, fast, and indispensable in real workflows.ConcretelyDesign and evolve our LLM / agent architectureOwn output quality across key use cases (emails, document analysis, etc.)Build evaluation systems (datasets, metrics, regression detection)Drive fast iteration loops from production dataImprove retrieval, reasoning, and tool usageEnsure production reliability (latency, failure modes, fallback)Work directly with product + founders on what to build and whyWhat This Role Is Really AboutMost teams fail because:they don’t know what “good output” meansthey don’t have evalsthey iterate randomlythey overuse agentsYour job is to fix that.You Will Turnvague user problems→ into structured AI systems→ with measurable performance→ that improve every weekWhat You Need To Be Excellent At Shipping real LLM systemsYou’ve built systems used in production (not demos)You understand RAG, tools, agents, structured outputsYou can design full pipelines, not just prompts Evaluation-driven developmentYou know how to define quality metricsYou build datasets from real usageYou run continuous evals to prevent regressions Debugging complex failuresYou can trace issues across:retrievalpromptsmodel behaviorYou don’t guess — you isolate and fix Speed of iterationYou move from problem → improvement in hours or days, not weeksYou use logs, traces, and data — not intuition alone Strong judgmentYou know when to:use an agent vs a pipelineadd complexity vs simplifyYou optimize for reliability and user value, not noveltyWhat We Don’t Care AboutNumber of years of experienceWhether you’ve used a specific frameworkFancy research credentialsIf you can build, debug, and improve real systems, you’re a fit.What Success Looks Like (first 90 Days)Clear eval framework for core use casesMeasurable improvement in output qualityFaster iteration cycles across the teamReduced hallucinations / failuresStronger system architecture decisionsStack (context, Not Requirements)Python (FastAPI)PostgresGoogle CloudLangGraph / LangChain (evolving)PostHog (product analytics)Langfuse (LLM traces)LLM APIs (Azure OpenAI)Sourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.