R

Manager, Datacenter Network Engineering

Runpod
11 days ago
Full-time
Remote
Worldwide
Remote Engineering

Runpod is pioneering the future of AI and machine learning, offering cutting-edge cloud infrastructure for full‑stack AI applications. Founded in 2022, we are a rapidly growing, well‑funded, remote‑first company with a global team across the US, Canada, and Europe. Our mission is to create a foundational platform that enables developers and companies to build, deploy, and scale custom AI systems with speed and flexibility.

We are looking for an Engineering Manager, Datacenter Network Engineering to lead the team responsible for designing, deploying, and operating Runpod’s global datacenter and backbone network. This role manages engineers working on L2/L3 fabrics, high-performance GPU networking, and global WAN connectivity that underpin our AI platform.

You will lead execution across multiple regions and vendors, while setting technical direction for network architecture that supports massive east-west traffic, low-latency GPU collectives, and secure multi-tenant isolation. This role is hands-on at the architectural level, while focused on team leadership, operational excellence, and scalability.

Responsibilities

  • Lead the Datacenter Networking Team: Manage and grow a team of network engineers responsible for datacenter fabrics, interconnects, and global WAN connectivity. Provide mentorship, technical guidance, and clear ownership boundaries.
  • Own Datacenter Network Architecture: Define and evolve network designs for GPU-heavy clusters, including spine-leaf topologies, ECMP routing, and high-bandwidth east-west traffic patterns.
  • High-Performance GPU Networking: Oversee design and operation of InfiniBand and RoCE-based fabrics supporting distributed training and inference workloads. Ensure performance, loss characteristics, and congestion control meet AI workload requirements.
  • Encapsulation & Overlay Protocols: Guide implementation and operations of encapsulation technologies such as VXLAN, EVPN, Geneve, or similar, enabling scalable multi-tenant isolation and flexible network provisioning.
  • Global WAN & Backbone Connectivity: Lead strategy and execution for global WAN connectivity, including private backbone links, IX connectivity, and hybrid connectivity with cloud providers and partners.
  • Reliability & Operations: Establish operational best practices for monitoring, capacity planning, change management, incident response, and post-mortems across the network stack.
  • Cross-Functional Collaboration: