🇪🇸 Spain · 7h ago

Security Software Engineer - Python & AI Evaluation

Braintrust

LinkedInleadEnglish-friendly
Help a top AI lab evaluate and improve large language models through security-focused coding tasks. Bring your software-engineering judgment and hands-on security experience to work involving vulnerabilities, exploit verification and security patches.This is a contracting engagement, with potential for a longer-term engagement. Remote candidates in the selected countries are elegible.What you will doEvaluate coding tasks involving software vulnerabilities, exploit verification and security patches.Create high-quality coding prompts and reference answers for benchmark-style problems.Evaluate model outputs for code generation, refactoring, debugging and implementation.Identify and document model failures, edge cases and reasoning gaps.Compare private language models with leading external models.Build or configure coding environments for evaluation and reinforcement learning.Follow detailed annotation and evaluation guidelines consistently.What you bringAt least five years of professional software-development experience and strong Python skills.Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting or validating security patches.The ability to apply structured evaluation criteria and write clear technical feedback.Fluency in written and spoken English.Helpful, not requiredProfessional code review, coding annotation, LLM/code evaluation or benchmark design.Knowledge of another programming language.Team leadership or mentoring experience.

Sourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.