VFUSION AI / TECHNOLOGY

The whole chain. Made visible.

From document to answer, GPU to API. This is our reference architecture: components are selected for each engagement, with explicit boundaries and measurement points.

VF / 01REFERENCE ARCHITECTURE
YOUR ENVIRONMENT01People and applicationsCHAT / API / DEVELOPER TOOLS02Access · API gatewayIDENTITY / LIMITS / ROUTING03AI agents, retrieval and modelsTOOLS / SOURCES / MODELS04GPU · Storage · MLOpsKUBERNETES / NVMe-oF / OBSERVABILITYON-PREMISES / EUROPEAN-HOSTED / AIR-GAPPED
One controlled chain. One clear boundary.
  1. People and applications: CHAT / API / DEVELOPER TOOLS
  2. Access · API gateway: IDENTITY / LIMITS / ROUTING
  3. AI agents, retrieval and models: TOOLS / SOURCES / MODELS
  4. GPU · Storage · MLOps: KUBERNETES / NVMe-oF / OBSERVABILITY

INSIDE THE BOUNDARY

Every layer has a purpose.

This diagram is a reference design. Exact components, active layers and access boundaries are defined for each project.

Open components. One coherent system.

Kubernetes, Kubeflow, Argo Workflows, Harbor and MLflow form the platform layer. Model serving and retrieval connect through controlled APIs.

Document processing as a foundation

Docling on GPU, extraction models for OCR and tables, and Gemma for visual interpretation. Dutch and English documents both matter.

Observe, test, operate

OpenTelemetry and Loki make processing stages visible. Quality evaluation, versioning and access controls belong in the architecture.

Agentic processing across applications

A shared AI interface connects model serving, document processing and orchestration tools to websites and apps. Data scope, allowed actions, approval and audit remain explicit for each application. Financial amounts and state transitions are validated by the application.

WHAT YOU GET

01

Language and code

gpt-oss, Gemma and an in-house quantised 72B coding model.

02

Retrieval

BGE-M3, Qwen3 embeddings, BM25 and reranking.

03

Vision and documents

YOLO, vision-language models, Docling and OCR.

04

Compute and fabric

H100, Kubernetes, Slurm, NVMe-oF and RoCE.

LAYER 01 → 10

RAG in ten layers.

A reference for design and evaluation. No promise of infallible answers; every layer has a defined responsibility.

01Ingest and normalise

Docling, OCR and orchestration turn source documents into consistent text and metadata.

02Two retrieval tracks

BM25 and dense embeddings combine exact search with semantic meaning.

03Select and rerank

HNSW search and reranking select the most relevant passages.

04Confidence gate

Relevance, freshness, authority and agreement determine whether there is enough evidence.

05Answer from context

The model works with selected sources and a versioned instruction.

06Cite the source

Claims connect to passages, documents and pages.

07Check grounding

An additional NLI check can flag missing or weak support.

08Measure quality

Evaluation sets, MLflow and alerts make changes visible.

09Cache with a policy

Exact and semantic caches with versions and expiry.

10Trace the chain

Per-stage tracing with OpenTelemetry and Loki.

THE NEXT STEP

What needs to work for you?

Start with the documents, the process or the technical question. We will help define a useful first step.

Book a conversation