The challenge
Give developers access to capable models on private infrastructure within a practical memory and capacity budget.
Dutch critical-infrastructure organisation
In-house quantisation reduced a 72B coding model from around 144 GB to 66 GB, fitting one H100 with a 90K-token context window.
Give developers access to capable models on private infrastructure within a practical memory and capacity budget.
In-house GPTQ quantisation, private model serving, an API gateway with keys and usage limits, and a shared chat interface. Embeddings, reranking and multimodal models fit the same access layer.
The memory reduction describes this implementation. It is not a general quality claim against larger or external models.
REFERENCE, NOT A NETWORK MAP
A generic overview of the building blocks. Technical addresses, internal systems and client identities are excluded from publication.
Explore the architecture ↗THE NEXT STEP
Start with the documents, the process or the technical question. We will help define a useful first step.