Dutch public-sector organisation

GPU capacity that follows the workload.

A GPU cluster where compute can be assigned to hosts through a PCIe fabric, with Slurm for workload scheduling.

The task

Use GPU capacity more flexibly than a fixed server-to-accelerator pairing allows.

What was built

A composable setup with RTX 8000 GPUs, a PCIe fabric and Slurm.

The scope

The project describes infrastructure and GPU assignment. It makes no unsubstantiated utilisation, cost or speed claims.

VF / 01REFERENCE ARCHITECTURE
YOUR ENVIRONMENT01People and applicationsCHAT / API / DEVELOPER TOOLS02Access · API gatewayIDENTITY / LIMITS / ROUTING03AI agents, retrieval and modelsTOOLS / SOURCES / MODELS04GPU · Storage · MLOpsKUBERNETES / NVMe-oF / OBSERVABILITYON-PREMISES / EUROPEAN-HOSTED / AIR-GAPPED
One controlled chain. One clear boundary.
  1. People and applications: CHAT / API / DEVELOPER TOOLS
  2. Access · API gateway: IDENTITY / LIMITS / ROUTING
  3. AI agents, retrieval and models: TOOLS / SOURCES / MODELS
  4. GPU · Storage · MLOps: KUBERNETES / NVMe-oF / OBSERVABILITY

REFERENCE, NOT A NETWORK MAP

The pattern behind the solution.

A generic overview of the building blocks. Technical addresses, internal systems and client identities are excluded from publication.

Explore the architecture ↗

THE NEXT STEP

What needs to work for you?

Start with the documents, the process or the technical question. We will help define a useful first step.

Book a conversation