🇩🇪 Berlin, Germany · 11h ago
Forward Deployed Engineer AI Inference
Lyceum
LinkedInEnglish-friendly
About LyceumLyceum is a sovereign European AI inference provider. We run open-source models on our own GPU infrastructure, powered by 100% renewable energy, so teams can build with AI on their own terms – without giving up their data or getting locked into a single vendor. Backed by tier-1 investors, we're growing fast and our inference business is scaling strongly.The RoleAs Senior Forward Deployed Engineer, you own the technical side of our largest and most complex inference deals, from first discovery call to go-live. You're the trusted technical counterpart for our customers' CTOs and ML leads, you design the setups that run their models, and you work hand in hand with our commercial team to get deals closed.You'll be one of two senior technical owners of our inference deals. Together with our Inference Sales Engineering lead, you'll also turn what we learn in the field into a repeatable playbook and product.What You'll DoOwn the technical side of our largest and most complex inference deals end to end, from discovery to go-liveBe the senior technical counterpart for customers' CTOs and ML leads on architecture, sizing, SLAs, latency/throughput trade-offs and costDesign dedicated inference setups: model choice, GPU type and count, parallelism, quantization, inference engine and configurationLead benchmarks and proofs of concept, and turn the results into clear recommendationsScope customer customizations with our engineering team and decide what becomes productBuild the playbook and tooling for matching workloads to GPUs, and feed our roadmap with what you learn in the fieldWork in tandem with our commercial team on pricing, proposals and the technical parts of contractsWhat We're Looking ForA degree in computer science, data science or a closely related field4+ years in a customer-facing technical role, e.g. solutions or sales engineer, forward deployed engineer, ML engineer working closely with customers, or technical consultantHands-on experience with LLM inference and model serving (e.g. vLLM, SGLang, TensorRT-LLM, Triton) and GPU sizingA track record of owning the technical side of complex B2B deals, ideally with enterprise customersYou're credible with customer CTOs and engineers alike, and you explain trade-offs clearlyAn entrepreneurial mindset: give you an outcome, and you find a way without getting blockedFluent EnglishBonus PointsExperience at an inference provider, GPU cloud or AI infrastructure companyBenchmarking and performance optimization: throughput, latency, cost per tokenExperience with quantization, parallelism strategies, KV-cache and batchingYou've coached or mentored juniorsStartup experienceGermanWhy Join UsReal ownership: Run the largest inference deals in our pipeline from day oneShape the playbook: Define how Lyceum sells and delivers dedicated inference, as one of two senior technical ownersProduct impact: Your field learnings become our matching and customization productFounder access: Work directly with our commercial leads and the foundersMission-driven team: Build sustainable, 100% renewable compute for the AI era, backed by top-tier investorsSourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.