A specialist field engineering service that brings high-density AI compute clusters from rack install to stable production workloads.
Added Jul 7, 2026
AI infrastructure buyers are deploying GPU?-dense clusters with complex dependencies across compute, storage, networking, cooling, firmware, Kubernetes, observability, and workload scheduling. The job signals show repeated hiring for people who can bridge customer architecture, data center operations, hardware bring-up, deployment automation, troubleshooting, and production support. This is not just a software problem; buyers need experienced operators who can make expensive clusters usable, reliable, and supportable at launch.
Start as a productized field engineering and commissioning service for neoclouds, colocation providers, sovereign AI programs, and enterprises standing up GPU? clusters. The first offer is a fixed-scope readiness and bring-up package covering facility checks, rack/network validation, firmware and driver baselining, Kubernetes or Slurm deployment review, GPU? burn-in, observability setup, runbooks, and handoff training. Over time, repeated checklists, telemetry collectors, reference architectures, and incident playbooks can become a repeatable managed service or lightweight software-assisted operating system for cluster readiness.
AI compute demand is pushing many buyers into unfamiliar high-density GPU? infrastructure faster than they can hire experienced deployment and operations teams. New NVIDIA-class platforms, liquid cooling, RoCE/InfiniBand fabrics, and hybrid cloud deployments create expensive failure modes during commissioning and early production.
Showing 1-20 of 20 signals
Monitor GPU cluster health and proactively communicate hardware issues to customers (thermal throttling, BMC failures, missing GPUs, and NVLink/InfiniBand degradation) with clear remediation steps Operate and maintain production infrastructure for enterprise GPU customers, including fleet rebalancing, Slurm cluster maintenance, node repair/migration, and Kubernetes-based workload management
You will work at the intersection of AI infrastructure, large-scale server deployment, data center engineering, and intelligent operations. This role offers the opportunity to participate in the deployment of next-generation AI/GPU clusters that power large-scale cloud and AI services. Rather than focusing on a single discipline, you will collaborate with hardware, networking, facilities, supply chain, software, and operations teams to deliver thousands of servers into production efficiently and
This role acts as the technical and operational interface between the customer/platform and Nebius infrastructure teams, ensuring reliable service delivery, SLA compliance, and smooth operations of large-scale GPU clusters and bare metal environments. You will coordinate across data center operations, network, hardware lifecycle, and infrastructure engineering teams to deliver world-class infrastructure services for large AI workloads.
We are looking for an experienced HPC Systems Engineer to support and operate large-scale Linux-based High-Performance Computing (HPC) environments. This role focuses on maintaining reliable, secure, and high-performance computing platforms that support research, academic, and enterprise workloads.
About the Role We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate.
+17 more signals