Dutch critical-infrastructure organisation

A 72B coding model on one H100.

In-house quantisation reduced a 72B coding model from around 144 GB to 66 GB, fitting one H100 with a 90K-token context window.

The challenge

Give developers access to capable models on private infrastructure within a practical memory and capacity budget.

What was built

In-house GPTQ quantisation, private model serving, an API gateway with keys and usage limits, and a shared chat interface. Embeddings, reranking and multimodal models fit the same access layer.

What the figures mean

The memory reduction describes this implementation. It is not a general quality claim against larger or external models.

VF / 01REFERENCE ARCHITECTURE
YOUR ENVIRONMENT01People and applicationsCHAT / API / DEVELOPER TOOLS02Access · API gatewayIDENTITY / LIMITS / ROUTING03AI agents, retrieval and modelsTOOLS / SOURCES / MODELS04GPU · Storage · MLOpsKUBERNETES / NVMe-oF / OBSERVABILITYON-PREMISES / EUROPEAN-HOSTED / AIR-GAPPED
One controlled chain. One clear boundary.
  1. People and applications: CHAT / API / DEVELOPER TOOLS
  2. Access · API gateway: IDENTITY / LIMITS / ROUTING
  3. AI agents, retrieval and models: TOOLS / SOURCES / MODELS
  4. GPU · Storage · MLOps: KUBERNETES / NVMe-oF / OBSERVABILITY

REFERENCE, NOT A NETWORK MAP

The pattern behind the solution.

A generic overview of the building blocks. Technical addresses, internal systems and client identities are excluded from publication.

Explore the architecture ↗

THE NEXT STEP

What needs to work for you?

Start with the documents, the process or the technical question. We will help define a useful first step.

Book a conversation