Jobgether
Jobgether

Staff Network Engineer (AI Fabric, Datacenter and Edge Networking)

engineeringfull-timeUK
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

Accountabilities

    • Own the long-term technical direction, architecture, and operational strategy for high-performance AI networking and broader infrastructure networking.
    • Design, deploy, and operate GPU networking fabrics optimized for distributed AI training and inference workloads.
    • Architect large-scale RoCE and Ethernet fabrics, including leaf-spine, fat-tree, rail, and multi-plane topologies.
    • Optimize east-west networking for GPU communication patterns and high-throughput distributed workloads, working closely with compute and platform teams.
    • Implement and operate technologies including RDMA, RoCE, high-bandwidth Ethernet, Spectrum-X, and related AI networking platforms.
    • Define reusable reference architectures, design principles, deployment standards, validation criteria, and operational patterns for future AI fabric deployments.
    • Evaluate architectural trade-offs across performance, resilience, scalability, cost, operability, and deployment speed, providing clear recommendations to stakeholders.
    • Design and operate Layer 2 and Layer 3 datacenter networks using technologies such as BGP, ECMP, EVPN/VXLAN, VLAN, VRF, OVS/OVN, and Linux networking.
    • Build scalable routing, segmentation, overlay, tenant-isolation, and north-south traffic-management architectures.
    • Design and maintain edge security and connectivity infrastructure, including WAF, TLS termination, DDoS mitigation, API gateways, and proxy protections.
    • Design and operate private inter-datacenter connectivity, including dark fiber, metro fiber rings, DWDM transport, high-capacity WAN links, and redundant backbone architectures.
    • Develop scalable principles for inter-site routing, redundancy, failure-domain isolation, and backbone evolution.
    • Lead end-to-end delivery of networking initiatives, from architecture and lab validation through production deployment and operational handover.
    • Collaborate with procurement and deployment teams on network bills of materials, datacenter layouts, rack elevations, capacity planning, performance modeling, and scaling strategies.
    • Establish safe network-change practices, including deployment validation, rollback procedures, acceptance criteria, and minimal-impact production execution.
    • Drive automation for provisioning, configuration management, monitoring, validation, and network lifecycle management using software engineering practices.
    • Own network reliability and operational performance, defining and tracking SLAs, SLOs, latency, recovery, and other key reliability metrics.
    • Serve as the senior escalation point for complex networking incidents, leading deep investigations, root-cause analysis, remediation, and long-term systemic improvements.
    • Build network observability and operational tooling that improves visibility, reliability, and day-two operations.
    • Act as the primary networking design authority, influencing platform architecture and helping adjacent teams understand how networking capabilities and constraints affect their decisions.
    • Mentor engineers across adjacent domains and help establish the standards, practices, and technical foundations for the future networking organization.
    • Requirements:

      • Extensive hands-on experience designing, deploying, and operating large-scale datacenter networks in production environments.
      • Expert knowledge of modern networking protocols and architectures, including BGP, OSPF, ECMP, and EVPN/VXLAN.
      • Proven experience operating high-speed Ethernet networks and networking hardware in production.
      • Strong experience with NVIDIA/Mellanox networking platforms and high-performance interconnect technologies.
      • Deep expertise designing and operating networking fabrics for large-scale GPU clusters and distributed AI workloads.
      • Strong understanding of GPU communication patterns and how distributed training and inference requirements influence network architecture and performance.
      • Practical experience with NCCL communication patterns, including all-reduce, all-gather, broadcast, and reduce-scatter, and their impact on network traffic.
      • Hands-on experience designing and tuning RoCE/RDMA fabrics for GPU clusters.
      • Strong understanding of RDMA transport behavior, failure modes, congestion, and performance characteristics.
      • Practical experience implementing and tuning congestion-management technologies such as PFC and ECN.
      • Experience designing rail-optimized GPU networking fabrics and diagnosing issues such as NCCL stalls, RDMA congestion, fabric hotspots, and packet loss affecting distributed workloads.
      • Understanding of how networking performance affects distributed AI frameworks such as PyTorch and TensorFlow.
      • Strong ability to troubleshoot complex cross-layer issues spanning hardware, firmware, operating-system networking, kernels, and distributed application communication.
      • Strong knowledge of networking hardware, optics, and high-speed interconnects, including dark fiber, DWDM, and high-capacity optical networking.
      • Experience designing network observability solutions and operational tooling.
      • Strong automation skills using Python, Bash, or similar technologies, with experience applying software engineering principles to infrastructure automation.
      • Ability to build reusable tools, standards, and validation approaches that increase engineering leverage across teams.
      • Proven ability to lead complex technical initiatives across engineering, operations, vendors, and other stakeholders.
      • Demonstrated ability to establish architectural direction and drive adoption of engineering standards through technical influence rather than formal authority.
      • Strong systems-level thinking, balancing performance, reliability, scalability, operational simplicity, and cost efficiency.
      • Strong communication skills, with the ability to clearly explain architectural decisions, risks, trade-offs, and technical recommendations.
      • Strong mentoring skills and willingness to raise the technical capabilities of engineers in adjacent domains.
      • Experience owning both architecture and direct implementation in a lean, rapidly scaling organization is strongly preferred.
      • Comfortable balancing immediate execution requirements with long-term architectural sustainability and operational maturity.
      • Willingness to travel as required across local, EMEA, US, APAC, or global locations.
      • Benefits:

        • Attractive compensation package reflecting your experience, technical expertise, and impact.
        • Flexible and hybrid-friendly working environment.
        • Opportunity to work within an internationally diverse and collaborative team.
        • Career growth opportunities within a fast-growing technology scale-up.
        • Significant autonomy and technical ownership over foundational networking architecture.
        • Opportunity to work on advanced AI infrastructure, GPU fabrics, high-speed networking, and distributed computing environments.
        • Opportunity to shape networking standards, automation, and operational practices from an early stage of the function.
        • Inclusive workplace committed to equal opportunity and diverse perspectives.
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now