A productized service that designs and operates the internal tooling layer for AI cloud and GPU? data center teams managing fast-growing compute fleets.
Added Jul 6, 2026
AI cloud providers and infrastructure-heavy AI companies are hiring senior platform, storage, network, and compute engineers to build the same internal operating layer: asset truth, hardware lifecycle workflows, monitoring, provisioning, performance testing, and operational automation. Existing DCIM and cloud tools do not fully cover GPU?-specific fleet realities such as rack-scale health, scheduler handoff, storage/network dependencies, and customer-facing capacity pressure. Teams end up stitching together CMDBs, spreadsheets, Terraform, Kubernetes, Slurm, observability systems, and custom scripts while growth outpaces operations.
Start as a high-trust implementation service for GPU? cloud operators: map their current deploy-to-operate workflow, clean up asset and topology data, integrate DCIM/CMDB with monitoring and provisioning systems, and build the first reliable operational workflows. The first deliverable can be a lightweight control plane for rack onboarding, hardware health, performance validation, incident handoff, and capacity visibility. Over time, repeated components become a productized toolkit for GPU? fleet lifecycle management rather than a generic platform SaaS?.
AI infrastructure buildout is forcing smaller GPU? clouds, enterprise AI labs, and specialized data center operators to manage fleets before they have mature internal platform teams. Job signals show companies repeatedly hiring expensive senior engineers for the same operational tooling gap.
Showing 1-20 of 20 signals
Monitor GPU cluster health and proactively communicate hardware issues to customers (thermal throttling, BMC failures, missing GPUs, and NVLink/InfiniBand degradation) with clear remediation steps Operate and maintain production infrastructure for enterprise GPU customers, including fleet rebalancing, Slurm cluster maintenance, node repair/migration, and Kubernetes-based workload management
Our mission is simple: build AI infrastructure that largely runs itself—where intelligent systems deploy, monitor, diagnose, optimize, and heal GPU fleets at massive scale. Every system you build will directly improve the speed, efficiency, and reliability of one of the world’s most advanced AI compute platforms.
Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators.
- End-to-End Production Readiness: Define launch criteria and readiness plans spanning server hardware, firmware, BMC, operating systems, drivers, GPU software stacks, networking, storage, security, telemetry, and operational tooling. - Rack-Scale Integration and Fleet Operations: Partner across hardware, data center, network, storage, power, cooling, and vendor teams to resolve system-level challenges and improve GPU fleet availability, utilization, serviceability, and lifecycle management.
As a Senior Software Engineer on the Automation team, you will design, build, and operate the services, APIs, and libraries that sit behind our OS image, payload, and boot-configuration systems — the software platform other HAVOCK engineers and partner teams rely on to release, test, and ship node software quickly and safely. You'll work on a constraint-solver–based service that resolves compatibility between images, kernels, drivers, payloads, and hardware into a single validated configuration; an end-to-end test framework that validates OS images on real hardware; a library suite for declaratively configuring node storage; and natural-language tooling that lets stakeholders query and interact with our systems. This is a software- and platform-engineering role first, with a clear forward trajectory toward AI-assisted automation — log triage, regression detection, natural-language interfaces to infrastructure — but the core of the job is designing and shipping reliable services and APIs.
+17 more signals