AI Cluster Fabric Planner
19 Signals+1

AI Cluster Fabric Planner

A planning and simulation tool for choosing GPU cluster network topology, optics, and power tradeoffs before procurement.

Added Jun 18, 2026

AI infrastructure
data center networking
procurement software
Opportunity score

Medium opportunity (54%)

The Problem

AI infrastructure buyers are no longer bottlenecked only by GPU count; cluster performance now depends heavily on backend fabric design, optics power draw, flow control, topology, and interconnect compatibility. Hyperscalers are using divergent approaches such as InfiniBand, custom Ethernet, optical circuit switching, SRD-style transport, and scale-up links, making vendor comparison difficult. Network architects need a concrete way to estimate whether a proposed cluster fabric will support training workloads without excess latency, congestion, power waste, or tenant-isolation risk.

Potential Solution

Build a SaaS tool where AI infrastructure teams model a planned GPU cluster by entering GPU type, rack count, switch family, optics type, topology, RDMA protocol, and expected collective communication pattern. The product produces topology diagrams, estimated fabric bandwidth, transceiver power load, congestion risk, failure-domain analysis, and a vendor-neutral comparison of Ethernet, InfiniBand, and emerging optical approaches. The first version can focus on procurement-stage planning using public vendor specs, user-entered BOMs, and simple all-reduce traffic models.

Why Now?

AI clusters are scaling past tens of thousands of GPUs while networking and optics consume a growing share of power and cost. New products such as 800G/1.6T switching, co-packaged optics, rail-optimized fabrics, and open accelerator interconnects make architecture decisions more complex and expensive to reverse.

Market validation
Search demand

Trend snapshot pending

Competition (0)

No matched competitors yet

Showing 1-19 of 19 signals

Job adsSep 3, 2026
nscale
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Drive cluster optimization end to end, tuning fabric configuration (adaptive routing, QoS and congestion control, SHARP in-network reduction, rail and topology-aware placement), collective communication libraries and algorithm selection, GPUDirect RDMA and storage paths, and host-level settings (huge pages, IRQ affinity, CPU governors, MIG and driver configuration) to convert raw hardware into delivered throughput. Partner with Infrastructure, Platform, SRE, and customer-facing teams to translat

Job adsAug 19, 2026
coreweave
Principal Solution Specialist, Core Services

Design the commercial framework for large-scale GPU reservation deals, including MFU modeling, cluster sizing, and network bandwidth commitments that support large enterprise closings. Partner with capacity, infrastructure, and networking teams to maintain a competitive edge on compute density, interconnect performance, and platform reliability across active and prospective customer deployments.

Job adsAug 11, 2026
meta
Telecom Conveyance Engineer, Data Center Infrastructure

Establish pathway capacity planning frameworks that account for current fill ratios, future growth, and evolving cable densities driven by AI/GPU infrastructure Create design automation tools, templates, and parametric models that enable scalable pathway design across multiple facility types and regions

Unlock 16 more signals

Go beyond the grade and inspect the evidence behind this opportunity.

Job ads

See which companies and roles are investing in this problem.
10 more

Podcast evidence

Read the exact transcript passages behind the idea.
5 more

Google Trends

Explore search interest, history, and momentum over time.
1 more