The experiment
591,000 training pairs, eight H100 GPUs and NeMo. Nine iterations led to a final training run of 32 hours.
In-house R&D
In-house research into adapting a 120B mixture-of-experts model. Training was completed; merge and export are not yet complete.
591,000 training pairs, eight H100 GPUs and NeMo. Nine iterations led to a final training run of 32 hours.
Attention-only LoRA to focus adaptation within the available compute and memory capacity.
A trained model is not yet a production service. Merge and export remain open; this case claims neither production deployment nor a general performance gain.
THE NEXT STEP
Start with the documents, the process or the technical question. We will help define a useful first step.