In-house R&D

120 billion parameters. Domain-specific training.

In-house research into adapting a 120B mixture-of-experts model. Training was completed; merge and export are not yet complete.

The experiment

591,000 training pairs, eight H100 GPUs and NeMo. Nine iterations led to a final training run of 32 hours.

The choice

Attention-only LoRA to focus adaptation within the available compute and memory capacity.

The current boundary

A trained model is not yet a production service. Merge and export remain open; this case claims neither production deployment nor a general performance gain.

THE NEXT STEP

What needs to work for you?

Start with the documents, the process or the technical question. We will help define a useful first step.

Book a conversation