🇵🇱 Warsaw, Poland · 21h ago
Kafka Reliability Engineer
Be | Shaping the Future Poland
LinkedInseniorEnglish-friendly
Be | Shaping the Future Poland has a proven position of being a reliable partner for financial services organisations to analyse complex requirements, find solutions and implement them in their entirety, regardless of their complexity. Since the foundation of Be Poland in 2013, we have been continually expanding and customising our spectrum of services. Today, we are privileged to have in our team the best individuals in each sector we operate within the financial services industry.Role: Senior SRE / Platform Engineer with KafkaLocation: fully remote from PolandContract Type: B2BWe are looking for experienced Senior SRE / Platform Engineers with Kafka to support a production readiness initiative for a critical booking processing platform. The assignment focuses on improving reliability, observability, resilience, and operational excellence across distributed systems and event-driven architectures. The ideal candidate combines strong hands-on experience with Kafka-based systems, observability tooling, cloud-native deployments, and Site Reliability Engineering practices.Key ResponsibilitiesObservability & MonitoringDefine and implement monitoring strategies based on Golden SignalsDesign and maintain dashboards, metrics, alerts, and reportingImprove centralized logging and distributed tracing capabilitiesDevelop alerting rules, thresholds, and operational runbooksSupport on-call processes and incident response activitiesKafka Reliability & MessagingDesign and optimize Kafka topics, partitioning strategies, and consumer groupsImplement retry mechanisms, dead-letter queues (DLQ), and idempotent processingDefine schema governance and messaging standardsMonitor Kafka performance, consumer lag, and throughputImprove reliability of asynchronous workflows, including backpressure handling and failure recoveryReliability Engineering & SRE PracticesDefine and manage SLIs, SLOs, and error budgetsParticipate in incident management and post-mortem activitiesDrive reliability-by-design principles across services and platformsIdentify and implement improvements that reduce operational overheadRelease & DeploymentSupport CI/CD pipelines and deployment automationImplement quality gates and release management processesWork with Canary, Blue-Green, and rollback strategiesSupport feature flag frameworks and version managementCollaborate on load testing and production readiness assessmentsWork with container orchestration platforms such as KubernetesData Protection & Disaster RecoveryDefine backup and restore strategiesSupport disaster recovery planning and testingContribute to RPO/RTO definitions and operational proceduresParticipate in DR exercises and resilience testingRequired Skills & ExperienceStrong experience as an SRE, Platform Engineer, DevOps Engineer, or Reliability EngineerHands-on experience with Apache Kafka in production environmentsExperience with monitoring, alerting, logging, and observability platformsKnowledge of distributed systems and event-driven architecturesExperience with Kubernetes and containerized environmentsExperience with CI/CD pipelines and deployment automationUnderstanding of incident management and operational excellence practicesExperience with MongoDB, including replication and backup conceptsStrong troubleshooting and problem-solving skillsFluent EnglishNice to HaveExperience in financial services or other highly regulated environmentsExperience with distributed tracing solutionsKnowledge of cloud platforms (AWS, Azure, or GCP)Experience implementing SLO/SLI frameworksCertifications related to Kubernetes, cloud technologies, or SRE practicesGerman languageOur offer:Competitive remuneration on B2B contractAccess to Mindgram – mental health & well-being platformFree gym at Q22Personal development – internal online / onsite DevTalksReferral bonus programInternational environmentSourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.