🇳🇱 Netherlands · 19h ago
AI Engineer Intern
LangWatch
LinkedInjuniorVisa sponsorEnglish-friendly
(only apply when you're associated to an university or recently graduated) No visa-sponsorships, or In a nutshellLangWatch builds tools that help AI teams understand, debug, and improve their LLM products. We're growing fast and looking for an AI Engineer Intern to own our benchmarking work: measuring how models and agents actually perform, on real tasks, and publishing what we find.This is a paid internship (Amsterdam-based) for 6 months with a strong chance of conversion for outstanding performance.About LangWatchLangWatch is an LLMOps platform for teams building with large language models. We help companies understand how users engage with their LLM features, what's working, and where to improve, enabling faster iteration and better user experiences. Our platform makes it easier to monitor, evaluate, and optimize AI products, closing the gap between proof-of-concept and reliable production.Our core is open source, we're backed by great VCs, and thousands of developers use what we ship. Now we're opening a hands-on internship for someone who wants to answer the question every AI team is asking: which model, which prompt, which setup, and how do we know?What you will be working on New models ship every week and every vendor claims their own benchmark. Teams building real products still can't answer whether a switch would help them. That gap is your project.You'll work closely with our CTO and the engineering team. Expect real ownership, code review that makes you better, and results that get read outside the company.Benchmark design: Build task suites that reflect what our users actually do, including agentic tool use, structured output, retrieval, and long-context work, rather than what looks good on a leaderboard.Harness engineering: Build and maintain the infrastructure to run benchmarks reproducibly across model providers, track cost and latency alongside quality, and re-run everything when a new model drops.Evaluation methodology: Work on the hard part, which is scoring. LLM-as-judge calibration, inter-rater agreement, variance across runs, and knowing when a difference is real.Analysis: Turn raw runs into findings that survive scrutiny, with error bars and honest caveats.Publishing: Write up results as reports, posts, and open datasets or repos that the community can check and reproduce.Product feedback: Feed what you learn back into LangWatch's evaluators, our gateway's model routing, and the guidance we give customers on model selection.Publish technical results in the open and handle the feedback that comes with it.Who should applyYou're studying CS/AI(or self-taught with strong projects) and can commit 6 months (full-time preferred; part-time 24 to 32h/week possible).You can write code. Python is essential here, typescript if you want to join the rest of the dev-team.You have a statistical instinct: sample sizes, variance, and significance are not new words.You've built something with LLMs, even if it was small or broken.You write clearly, and you can explain a result to someone who will try to poke holes in it.You care about open source and about work that others can reproduce.You (will) have tangible work to show: repos, demos, articles, notebooks, or contributions.Bonus pointsExperience with eval frameworks, agent frameworks, or RAG systems.You've run experiments at any scale and had to defend the methodology.Background in ML/AI, or experience integrating LLMs into apps for others.You've published technical results online before.Compensation & logisticsPaid internship: depending on location, hours, and experience.Location: Amsterdam-Zuid office.Perks: Friday lunches & drinks.Future: High-performing interns are considered first for return offers and full-time roles in Engineering, Research, or Product.Team & ways of workingSmall, ambitious, and collaborative. We prototype, test, and ship. We love ideas, and love them more when they're in users' hands. We value time together but support focused remote work.Sourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.