🇳🇱 Amsterdam, Netherlands · 6h ago
Senior Site Reliability Engineering
Brookwood Recruitment Ltd
LinkedInseniorEnglish-friendly
Are you a passionate engineer dedicated to ensuring the high availability and performance of critical systems? We are seeking a talented Site Reliability Engineer I (SRE I) to be a key player in designing, building, and maintaining resilient, scalable, and efficient services that power our innovative technological landscape. In this impactful role, you will:Develop and deliver robust software solutions using relevant programming languages, applying your expertise in systems, services, and tools.Evaluate and design architecture solutions aligned with business needs, ensuring scalability and future growth.Take ownership of end-to-end system management, from deployment through ongoing operations, monitoring health and performance metrics.Lead incident response efforts, diagnosing root causes and implementing long-term solutions to boost reliability.Automate repetitive tasks and optimize system efficiency to minimize operational toil.Enhance observability systems by reviewing and improving performance monitoring and alerting.Collaborate across teams, providing architectural guidance and mentoring junior engineers to foster a culture of continuous improvement.Required skills:Extensive experience in building software applications, with a strong understanding of systems architecture, Infrastructure as Code (Terraform), and cloud platforms such as AWS.Proven ability to evaluate complex system designs and recommend scalable, resilient, and cost-effective solutions.Experience with event-driven architectures and data streaming technologies, including Apache Kafka and Confluent Cloud.Expertise in incident management, automation, CI/CD, and deployment practices, with experience using tools such as Harness.Strong understanding of system observability, monitoring, and performance management, with experience using Prometheus and Grafana.Critical thinking skills to identify, troubleshoot, and resolve complex technical issues across distributed systems and cloud environments.Excellent communication skills, capable of conveying complex technical concepts clearly and effectively to both technical and non-technical stakeholders.Demonstrated experience in coaching, mentoring, and supporting colleagues or stakeholders, helping to promote technical best practices and knowledge sharing.Nice to have skills:Additional experience with cloud native technologies, vendor evaluations, and advanced automation tools.Certifications related to cloud platforms, security, or DevOps practices.Familiarity with specific monitoring or observability tools and frameworks.Preferred education and experience:Master's degree or higher in Computer Science, Software Engineering, or a related field.5-8 years of relevant, hands-on experience in site reliability engineering, software development, or systems architecture.Other requirements:Ability to work in a dynamic environment that values diversity and inclusion.Commitment to ongoing learning and professional growth.Flexibility to collaborate with global teams and across various time zones as required.Ready to help shape the future of reliable, scalable systems? We invite enthusiastic, innovative candidates who are eager to make an impact to apply now and be part of our forward-thinking engineering community!Sourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.