🇵🇱 Poland · 6h ago
SRE Lead
Thrive IT Systems
LinkedInleadEnglish-friendly
We are seeking a highly skilled Software Engineering Site Reliability Engineering SRE Lead to drive the design development testing and operational excellence of enterprise technology solutions The ideal candidate will possess strong software engineering fundamentals deep Software Development Lifecycle SDLC expertise extensive test automation experience and a proven track record in building and operating highly reliable scalable and resilient platformsKey Responsibilities:Lead the design development and deployment of highquality software solutions following modern engineering best practicesDrive endtoend SDLC processes including requirements analysis design development testing deployment and production supportDevelop and implement automated testing frameworks and strategies to improve software quality reliability and release velocityDesign and maintain CICD pipelines to support continuous integration automated testing and continuous deliveryApply SRE principles to improve system reliability scalability availability and operational efficiencyEstablish and monitor SLIs SLOs and error budgets to ensure service performance and reliability objectives are metCollaborate with development infrastructure security and operations teams to streamline software delivery and production supportLead root cause analysis incident management and postincident reviews to drive continuous improvementChampion observability practices through monitoring logging ing and performance analysisMentor engineering teams on software engineering best practices test automation and reliability engineeringRequired QualificationsHands on years of experience in software engineering and enterprise application developmentStrong expertise in Software Development Lifecycle SDLC methodologies and engineering best practicesHandson experience designing and implementing test automation frameworks and strategiesProven experience with Site Reliability Engineering SRE production operations and platform reliabilityExperience with CICD pipelines DevOps practices and release automationStrong understanding of system performance scalability availability and resiliency principlesExperience with monitoring observability incident management and operational excellenceExcellent analytical troubleshooting and problemsolving skillsStrong communication and leadership abilitiesPreferred SkillsCloud platforms AWS Azure or Google Cloud PlatformContainerization and orchestration technologies such as Docker and KubernetesInfrastructure as Code Terraform Ansible or similarMonitoring and observability tools such as Prometheus Grafana Datadog Splunk or DynatraceAgile Scrum and DevSecOps experienceTechnical SkillsSoftware EngineeringSDLC ManagementTest AutomationSite Reliability Engineering SRECICD and DevOpsSystem Reliability Performance EngineeringObservability MonitoringIncident Problem ManagementCloud and Container TechnologiesAutomation Infrastructure as Code IaCSkillsMandatory Skills : Application Security (application security framework/ threat modelling/ Secure SDLC/ DevSecOps/Application Security Architecture Review)Sourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.