🇩🇪 Berlin, Germany · 1h ago
Member of Technical Staff (Search Crawler Analyst)
Perplexity
LinkedInEnglish-friendly
The internet is vast, containing trillions of URLs. Perplexity’s crawling and storage system is complex and has multiple stages (URL discovery, crawling, parsing, indexing). Each stage offers many opportunities for improvement and room for intricate bugs. You’ll work at the intersection of data analysis and engineering — designing metrics, building data pipelines, and improving the quality of our search and answer systems.ResponsibilitiesFind and diagnose quality issues in our crawling pipelineTrain small models that optimize particular aspects of the pipeline (e.g. parsing quality)Build datasets for model training, including LLM-as-a-judge labeling pipelinesImprove page selection algorithms for indexingDesign and analyze experiments to validate improvementsQualifications4+ years of experience as a data analyst, ML engineer, or in a related roleStrong coding skills: you should be able to write production-grade code at the level of a mid-level backend engineerExperience designing metrics from scratchExperience training ML models that shipped to production with measurable metric improvementsNice to haveDirect experience working on web crawling or indexing pipelinesSourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.