J
Senior Site Reliability Engineer (SRE, Compute Node Team)
Jobgether
1 day ago
Full-time
Remote
Worldwide
Remote Engineering
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer (SRE, Compute Node Team) based in Germany. This is a senior Site Reliability Engineering role focused on the infrastructure that runs and manages virtual machines across a large-scale cloud platform.You will work close to the Linux operating system, hypervisor, and node-level services that form the foundation of the compute environment.The role combines deep Linux systems engineering, virtualization, containerization, observability, and production reliability.You will investigate complex issues involving CPU, memory, NUMA, cgroups, scheduling, and system performance across user and kernel space.You will also help shape reliability practices through strong monitoring, incident response, root-cause analysis, and postmortem processes.Collaboration with platform, kernel, hypervisor, GPU, and infrastructure teams will be central to improving system design and operability.This is an opportunity to influence critical compute infrastructure supporting demanding AI and cloud workloads at significant scale.