Jobgether
Staff Network Engineer (AI Fabric, Datacenter and Edge Networking)
engineeringfull-timeUK
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more
About the role
Accountabilities
- Own the long-term technical direction, architecture, and operational strategy for high-performance AI networking and broader infrastructure networking.
- Design, deploy, and operate GPU networking fabrics optimized for distributed AI training and inference workloads.
- Architect large-scale RoCE and Ethernet fabrics, including leaf-spine, fat-tree, rail, and multi-plane topologies.
- Optimize east-west networking for GPU communication patterns and high-throughput distributed workloads, working closely with compute and platform teams.
- Implement and operate technologies including RDMA, RoCE, high-bandwidth Ethernet, Spectrum-X, and related AI networking platforms.
- Define reusable reference architectures, design principles, deployment standards, validation criteria, and operational patterns for future AI fabric deployments.
- Evaluate architectural trade-offs across performance, resilience, scalability, cost, operability, and deployment speed, providing clear recommendations to stakeholders.
- Design and operate Layer 2 and Layer 3 datacenter networks using technologies such as BGP, ECMP, EVPN/VXLAN, VLAN, VRF, OVS/OVN, and Linux networking.
- Build scalable routing, segmentation, overlay, tenant-isolation, and north-south traffic-management architectures.
- Design and maintain edge security and connectivity infrastructure, including WAF, TLS termination, DDoS mitigation, API gateways, and proxy protections.
- Design and operate private inter-datacenter connectivity, including dark fiber, metro fiber rings, DWDM transport, high-capacity WAN links, and redundant backbone architectures.
- Develop scalable principles for inter-site routing, redundancy, failure-domain isolation, and backbone evolution.
- Lead end-to-end delivery of networking initiatives, from architecture and lab validation through production deployment and operational handover.
- Collaborate with procurement and deployment teams on network bills of materials, datacenter layouts, rack elevations, capacity planning, performance modeling, and scaling strategies.
- Establish safe network-change practices, including deployment validation, rollback procedures, acceptance criteria, and minimal-impact production execution.
- Drive automation for provisioning, configuration management, monitoring, validation, and network lifecycle management using software engineering practices.
- Own network reliability and operational performance, defining and tracking SLAs, SLOs, latency, recovery, and other key reliability metrics.
- Serve as the senior escalation point for complex networking incidents, leading deep investigations, root-cause analysis, remediation, and long-term systemic improvements.
- Build network observability and operational tooling that improves visibility, reliability, and day-two operations.
- Act as the primary networking design authority, influencing platform architecture and helping adjacent teams understand how networking capabilities and constraints affect their decisions.
- Mentor engineers across adjacent domains and help establish the standards, practices, and technical foundations for the future networking organization.
- Extensive hands-on experience designing, deploying, and operating large-scale datacenter networks in production environments.
- Expert knowledge of modern networking protocols and architectures, including BGP, OSPF, ECMP, and EVPN/VXLAN.
- Proven experience operating high-speed Ethernet networks and networking hardware in production.
- Strong experience with NVIDIA/Mellanox networking platforms and high-performance interconnect technologies.
- Deep expertise designing and operating networking fabrics for large-scale GPU clusters and distributed AI workloads.
- Strong understanding of GPU communication patterns and how distributed training and inference requirements influence network architecture and performance.
- Practical experience with NCCL communication patterns, including all-reduce, all-gather, broadcast, and reduce-scatter, and their impact on network traffic.
- Hands-on experience designing and tuning RoCE/RDMA fabrics for GPU clusters.
- Strong understanding of RDMA transport behavior, failure modes, congestion, and performance characteristics.
- Practical experience implementing and tuning congestion-management technologies such as PFC and ECN.
- Experience designing rail-optimized GPU networking fabrics and diagnosing issues such as NCCL stalls, RDMA congestion, fabric hotspots, and packet loss affecting distributed workloads.
- Understanding of how networking performance affects distributed AI frameworks such as PyTorch and TensorFlow.
- Strong ability to troubleshoot complex cross-layer issues spanning hardware, firmware, operating-system networking, kernels, and distributed application communication.
- Strong knowledge of networking hardware, optics, and high-speed interconnects, including dark fiber, DWDM, and high-capacity optical networking.
- Experience designing network observability solutions and operational tooling.
- Strong automation skills using Python, Bash, or similar technologies, with experience applying software engineering principles to infrastructure automation.
- Ability to build reusable tools, standards, and validation approaches that increase engineering leverage across teams.
- Proven ability to lead complex technical initiatives across engineering, operations, vendors, and other stakeholders.
- Demonstrated ability to establish architectural direction and drive adoption of engineering standards through technical influence rather than formal authority.
- Strong systems-level thinking, balancing performance, reliability, scalability, operational simplicity, and cost efficiency.
- Strong communication skills, with the ability to clearly explain architectural decisions, risks, trade-offs, and technical recommendations.
- Strong mentoring skills and willingness to raise the technical capabilities of engineers in adjacent domains.
- Experience owning both architecture and direct implementation in a lean, rapidly scaling organization is strongly preferred.
- Comfortable balancing immediate execution requirements with long-term architectural sustainability and operational maturity.
- Willingness to travel as required across local, EMEA, US, APAC, or global locations.
- Attractive compensation package reflecting your experience, technical expertise, and impact.
- Flexible and hybrid-friendly working environment.
- Opportunity to work within an internationally diverse and collaborative team.
- Career growth opportunities within a fast-growing technology scale-up.
- Significant autonomy and technical ownership over foundational networking architecture.
- Opportunity to work on advanced AI infrastructure, GPU fabrics, high-speed networking, and distributed computing environments.
- Opportunity to shape networking standards, automation, and operational practices from an early stage of the function.
- Inclusive workplace committed to equal opportunity and diverse perspectives.
Requirements:
Benefits:
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply