Skip to content
AI Solutions

A model of your own, when you actually need one.

Most problems are solved by better retrieval, not a custom model. When one is genuinely warranted — a specialised domain, an offline requirement, or data that cannot leave your building — iLeaf fine-tunes and deploys models on your own infrastructure, including full on-premise GPU deployments in regulated environments.

Ask Lia about custom models

What this includes

Feasibility assessment
An honest read on whether a custom model beats retrieval or prompting for your case, before you fund one.
Fine-tuning
Supervised fine-tuning and adapters on open-weight models, sized to the task rather than the largest available.
Dataset engineering
Building, labelling and versioning the training set — usually the hard part and the one that decides the result.
On-premise deployment
Models served on your own GPUs where data cannot leave, with the same governance as any hosted deployment.
Classical ML
Gradient boosting and time-series where they outperform a language model, which for tabular prediction is often.
Computer vision
Detection, classification and OCR pipelines for image-based products and document capture.

How we run it

Every phase ships something usable on its own, so you are never holding a half-finished system waiting on the next milestone.

  1. Try the cheap thing first

    Prompting and retrieval are benchmarked before fine-tuning is considered. Often they win, and we will say so.

  2. Build the eval before the model

    A held-out set with agreed metrics comes first, so improvement is provable rather than a matter of impression.

  3. Start small

    The smallest model that clears the bar wins — it is cheaper to serve, faster to respond and easier to run on-premise.

  4. Plan for retraining

    Drift monitoring and a retraining path are part of the first delivery, not a later discovery.

What we build it with

  • PyTorch
  • Hugging Face
  • LoRA / PEFT
  • vLLM
  • Open-weight models
  • scikit-learn
  • XGBoost
  • ONNX
  • NVIDIA GPU
  • Weights & Biases
  • Airflow
  • Docker

Questions we get asked

Do we need a custom model?

Usually not. Retrieval over your own data, with good prompting, handles most enterprise cases at a fraction of the cost and with no retraining burden. A custom model earns its place when the domain language is genuinely specialised, latency or offline use rules out an API, or data cannot leave your infrastructure.

Can models run entirely inside our environment?

Yes. We have delivered clinical AI where every component, including the models, ran on the hospital's own GPUs with no data leaving the premises. Open-weight models make this practical, and the governance layer works identically whether the model is hosted or local.

Who owns the model and the training data?

You do — the weights, the dataset and the pipeline that produced them, in your own accounts. There is no arrangement where your data improves a model we then sell to someone else.

Let’s talk about custom models.

Tell us what you are running and what it needs to do next. We will tell you honestly whether we are the right team for it.

Talk to a solutions lead