🇪🇸 Spain · 3h ago
Senior Software Engineer (Capacity and Quota Management)
Jobgether
LinkedInseniorEnglish-friendly
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer (Capacity and Quota Management) based in Spain.Join a high-impact engineering team responsible for controlling how compute resources are allocated across a large-scale AI cloud platform.You will build and evolve the control plane that determines who receives compute resources, when they receive them, and how much capacity they can use.The role covers both GPU capacity reservations and quota management across multiple cloud services.You will tackle challenging distributed-systems problems involving resource allocation, consistency, reliability, rebalancing, and automated reclamation.Your work will directly influence infrastructure efficiency, customer experience, and the ability to deliver reliable compute at scale.This is a high-ownership position where you will help shape technical foundations, architectural decisions, and long-term platform evolution.You will work in a fast-moving, international engineering environment with significant autonomy and the opportunity to contribute to the future of AI infrastructure.AccountabilitiesDesign, develop, and evolve the Capacity and Quota Management control plane responsible for allocating and managing compute resources.Build and maintain a capacity reservation engine that manages guaranteed GPU reservations throughout their lifecycle.Develop and improve resource rebalancing and automated reclamation mechanisms to maximise infrastructure utilisation while maintaining customer commitments.Build and evolve the Quotas Control Plane responsible for managing quotas across multiple services and handling customer and internal quota requests.Design and implement resource distribution algorithms that balance capacity availability, customer requirements, and operational constraints.Develop reliable APIs that provide external customers, internal users, and downstream services with visibility and control over resource limits and allocations.Solve complex distributed-systems challenges involving transactions, consistency, asynchronous processing, and service reliability.Contribute to the technical architecture and long-term evolution of the platform, making decisions that remain effective beyond individual features or projects.Collaborate with engineers and stakeholders to translate infrastructure and product requirements into scalable, maintainable technical solutions.Take strong ownership of engineering outcomes, proactively identifying opportunities to improve system reliability, scalability, resource efficiency, and operational performance.Requirements5+ years of professional software engineering experience, ideally working on backend or distributed systems.Strong knowledge of Python, or the willingness and ability to become productive quickly with the relevant technology stack.Proven experience designing and building distributed backend systems and microservice architectures.Solid understanding of distributed transactions, consistency models, and reliable asynchronous workflows.Experience developing scalable, resilient, and maintainable services in production environments.Strong software engineering fundamentals and the ability to reason about complex system behaviour and architectural trade-offs.Strong communication skills and the ability to collaborate effectively with technical and cross-functional stakeholders.A strong ownership mindset, with the autonomy and judgment to make technical decisions and drive initiatives from concept through implementation.Ability to work effectively in a fast-moving, international environment where priorities and technical challenges evolve quickly.Experience with resource management, cloud infrastructure, capacity planning, quota systems, scheduling, or large-scale infrastructure platforms would be advantageous.BenefitsCompetitive compensation.Opportunities for career growth, continuous learning, and professional development.Flexible working arrangements with a high degree of autonomy and ownership.Opportunity to work on impactful AI and cloud infrastructure projects.Collaborative and innovative international working environment.Exposure to complex, large-scale distributed systems and challenging infrastructure problems.Meaningful opportunity to influence architecture, technical foundations, and the future evolution of the platform.Work alongside experienced engineers and talented teams operating across an international environment.How Jobgether WorksWe use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.We appreciate your interest and wish you the best! Why Apply Through Jobgether?Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.Sourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.