🇮🇹 Milan, Italy · 17h ago
Network Architect Engineer
Stealth Startup
LinkedInseniorEnglish-friendly
About UsWe are an early-stage technology company operating at the intersection of Energy and AI infrastructure, focused on developing a new generation of efficient and scalable computing infrastructure.Our approach combines distributed computing, modular infrastructure and energy efficiency to support the rapidly growing demand for AI and high-performance computing.The company is developing a new model for deploying compute capacity in a flexible and scalable way, with a strong focus on efficiency, sustainability, reliability and responsible infrastructure development.This is an opportunity to join a growing team at an early stage and contribute directly to the development of a new infrastructure platform, working at the intersection of power, data centers and advanced computing.Role OverviewWe are seeking a Network Architect Engineer to design, build and own the networking of our next-generation AI/HPC compute infrastructure, and of the meshed clusters these systems form. The role spans two connected worlds: the high-performance fabric inside the compute infrastructure (GPU/HPC interconnect, low-latency traffic) and the distributed network between geographically dispersed sites (a sovereign, resilient mesh across multiple locations). You will define the network architecture end to end, and you will be hands-on implementing, validating, and operating it. We do not expect deep mastery of both worlds on day one: we are looking for a strong architect who is excellent in at least one and has the appetite and ability to grow into the other.What You Will DeliverThis is a high-impact, high-ownership role. You will define how our compute infrastructure connects its GPUs, storage and management systems, how distributed sites communicate with each other, and how the infrastructure connects to the outside world. Your architecture will directly influence every system and cluster we deploy.In your first year: You will own the network design and architecture of our first-generation compute infrastructure, evolving it into a production-ready system. This starts with a working intra-system fabric (compute/GPU interconnect, storage and management/out-of-band networks) and an initial inter-site interconnect validated in the lab, driving technology selection (switching, NICs/DPUs, overlay and routing protocols) and working hands-on to bring the network to life in the physical infrastructure. From there, working closely with platform, hardware, and security colleagues, you will industrialize the fabric and validate a multi-site sovereign mesh, proven on real deployments, delivering a design that is automated (infrastructure-as-code and observable) and ready to scale into series deployment.Within 6 months: You will own and deliver the network design of our first compute infrastructure prototype: a working intra-system fabric (compute/GPU interconnect, storage and management/out-of-band networks) and the first inter-site interconnect validated in the lab. This means driving technology selection (switching, NICs/DPUs, overlay and routing protocols), making key design decisions, and working hands-on to bring the network to life in the physical infrastructure.Within 12 months: You will own the network architecture of our first production-ready system: an industrialized intra-system fabric and a validated multi-site sovereign mesh prototype, proven on real deployments. Working closely with platform, hardware, and security colleagues, you will deliver a design that is automated (infrastructure-as-code and observable), and ready to scale into series deployment.Key ResponsibilitiesDefine and document the end-to-end network architecture of the compute infrastructure and of the meshed cluster, across compute fabric, storage, management/out-of-band, and inter-site connectivityDesign and implement the intra-system high-performance fabric (high-speed Ethernet and/or RDMA transports such as RoCE/InfiniBand), including topology, NIC/DPU and switch selection, and congestion/QoS tuningDesign and implement the inter-site overlay: a resilient sovereign mesh (e.g. WireGuard/IPsec), dynamic routing (BGP/OSPF), multi-zone segmentation, and failover across heterogeneous and sometimes constrained uplinksBuild network automation and infrastructure-as-code (provisioning, configuration, and lifecycle management) so the network is reproducible across every system we deployStand up network observability: telemetry, monitoring, and troubleshooting workflows for both intra-system and inter-site trafficCollaborate with electrical/mechanical engineers on physical integration (cabling, connectors, port layout, serviceability) and with platform/security colleagues on segmentation and tenant isolationProduce network diagrams, IP plans, cabling and BOM inputs, and operational documentation to support fabrication, deployment, and field operationsSupport prototype builds, lab testing, and field validation; incorporate findings into design iterations and contribute to cross-functional design reviewsRequirementsBachelor’s degree in Electrical Engineering, Computer Science, Networking, or equivalent field4-10 years of hands-on networking experience in datacenter, HPC, service-provider, or distributed/edge environmentsDeep expertise in at least one of: (a) datacenter/HPC fabrics (high-speed Ethernet, RDMA/RoCE/InfiniBand, leaf-spine, low-latency east-west) or (b) distributed/WAN networking (overlay meshes, VPN, BGP, SD-WAN, multi-site resilience), with the appetite and ability to grow into the otherStrong command of IP networking fundamentals: routing and switching, VLAN/VXLAN, overlays, and network security conceptsProficiency in network automation and infrastructure-as-code (e.g. Python, Ansible, Terraform, NetBox or similar)Ability to work in a fast-paced, early-stage environment with a high degree of ownership and autonomyEnglish speakingNice to HaveHands-on experience with GPU/AI or HPC interconnects and RDMA tuning (RoCEv2, PFC/ECN/DCQCN, InfiniBand)Experience with SmartNICs/DPUs (e.g. NVIDIA BlueField) and open/whitebox networking (SONiC, FRR, Cumulus)Experience operating large-scale overlay meshes or SD-WAN across many sitesExperience with edge or outdoor deployments and constrained/heterogeneous connectivity (fiber, cellular, satellite backup)Familiarity with zero-trust networking and network segmentation for multi-tenant infrastructurePrior experience at a hardware or infrastructure startup or scale-up, where scope and pace require broad ownershipLocation Milan, Italy - hybridSourced from LinkedIn. Relocantly aggregates public job postings; apply on the original site.