A model of your own, when you actually need one.
Most problems are solved by better retrieval, not a custom model. When one is genuinely warranted — a specialised domain, an offline requirement, or data that cannot leave your building — iLeaf fine-tunes and deploys models on your own infrastructure, including full on-premise GPU deployments in regulated environments.
Ask Lia about custom modelsWhat this includes
- Feasibility assessment
- An honest read on whether a custom model beats retrieval or prompting for your case, before you fund one.
- Fine-tuning
- Supervised fine-tuning and adapters on open-weight models, sized to the task rather than the largest available.
- Dataset engineering
- Building, labelling and versioning the training set — usually the hard part and the one that decides the result.
- On-premise deployment
- Models served on your own GPUs where data cannot leave, with the same governance as any hosted deployment.
- Classical ML
- Gradient boosting and time-series where they outperform a language model, which for tabular prediction is often.
- Computer vision
- Detection, classification and OCR pipelines for image-based products and document capture.
How we run it
Every phase ships something usable on its own, so you are never holding a half-finished system waiting on the next milestone.
Try the cheap thing first
Prompting and retrieval are benchmarked before fine-tuning is considered. Often they win, and we will say so.
Build the eval before the model
A held-out set with agreed metrics comes first, so improvement is provable rather than a matter of impression.
Start small
The smallest model that clears the bar wins — it is cheaper to serve, faster to respond and easier to run on-premise.
Plan for retraining
Drift monitoring and a retraining path are part of the first delivery, not a later discovery.
What we build it with
- PyTorch
- Hugging Face
- LoRA / PEFT
- vLLM
- Open-weight models
- scikit-learn
- XGBoost
- ONNX
- NVIDIA GPU
- Weights & Biases
- Airflow
- Docker
Questions we get asked
Do we need a custom model?
Usually not. Retrieval over your own data, with good prompting, handles most enterprise cases at a fraction of the cost and with no retraining burden. A custom model earns its place when the domain language is genuinely specialised, latency or offline use rules out an API, or data cannot leave your infrastructure.
Can models run entirely inside our environment?
Yes. We have delivered clinical AI where every component, including the models, ran on the hospital's own GPUs with no data leaving the premises. Open-weight models make this practical, and the governance layer works identically whether the model is hosted or local.
Who owns the model and the training data?
You do — the weights, the dataset and the pipeline that produced them, in your own accounts. There is no arrangement where your data improves a model we then sell to someone else.
Let’s talk about custom models.
Tell us what you are running and what it needs to do next. We will tell you honestly whether we are the right team for it.
