You want AI on the documents and records that matter, but that data cannot go to a cloud service. The answer is an AI platform that runs on your own hardware, inside your own network, so the data never has to leave. We size it, specify it, install it and hand it over ready to use, then train your team to run it themselves.
Could your staff use AI on confidential documents and personal data tomorrow, without any of it leaving your network?
For most organizations the answer is no. Cloud AI services are ruled out for sensitive material, so the AI project stalls, or staff paste confidential text into public chatbots anyway. Buying a GPU machine does not fix that by itself. Someone still has to size it, build a platform that stays up, install models that work in Thai, and hand over something your own team can run.
This is not an argument for running everything locally. Public and low-sensitivity work can still go to the best model in the cloud. This platform is where the work that must stay inside runs. How that split can be decided per request, by the data's own classification, is set out in Your Best Model Is in the Cloud.
This is the AI domain's delivery engagement: the step that turns ‘we cannot use AI on this data’ into a working platform inside your walls.
Six deliverables. The hardware is yours, the platform is yours, and so is every password to it.
How large a model you need, how much GPU memory it takes, and how many people can ask questions at the same time. Plus the CPU, memory and storage for everything that runs around the model. Written down, with the working shown.
A specification more than one vendor can meet, written the way public procurement needs it. Buy through your own process or through us. The specification does not tie you to a brand.
A virtual machine platform on ZFS storage: mirrored system disks, redundant VM storage, snapshots, and backups kept both on the machine and off it, following the 3-2-1 rule. The GPU is passed through to the virtual machines that need it.
Model serving, models chosen to work in Thai, a chat interface for staff, a starter for asking questions of your own documents, and an API your developers can build on.
Separate environments from the start. When production moves to server-class hardware, the workstation becomes your development machine instead of something you replace.
A runbook written for your team, hands-on training, and support terms defined in writing: what is covered, when, and how quickly we respond.
A workstation-class machine is the right answer for development, and for pilot production where a short outage is acceptable. It is not the right answer for AI your organization cannot run without, around the clock. That belongs on server-class hardware, with redundant power and components.
We tell you which one you need before you buy, and the platform we build moves to the server when you get there.
The platform is built to grow. Each of these is a defined next step, and you are free to take any of them yourself or with anyone else.
Three questions decide the machine: how large a model the work needs, how much GPU memory that model takes, and how many people will ask questions at the same time. The model is loaded once. The memory left over is what serves the people asking, and each person asking needs their own share of it.
How many people can ask at once, by GPU memory on one card (planning rule: 2 GB per person asking)
People asking at the same moment (not total staff)
GPU memory on one card (GB)
The same figures, as a table
| Model size (4‑bit) | Model takes | 24 GB | 48 GB | 96 GB |
|---|---|---|---|---|
| ~12B | ~7 GB | 7 | 17 | 39 |
| ~70B | ~43 GB | — | — | 21 |
| ~120B | ~65 GB | — | — | 10 |
Assumes 4-bit models, about 2 GB of memory for each person asking, and 10% of the card kept free for the software itself. A dash means the model does not fit, or leaves no room for anyone to ask.
2 GB is a planning figure, deliberately on the safe side. The real amount depends on the model's design and how long each conversation is: from under 1 GB to about 3 GB per person for a typical conversation, and more for long documents. Memory is also only the ceiling. How fast answers come back with many people asking at once is the other limit, and we check both during sizing.
Automated jobs count too: every bot or scheduled task that asks the AI takes a place, the same as a person.
GPU memory decides how big a model you can run and how many can ask at once. CPU and RAM decide how much application you can build around it.
The scope bands on the chart are a rough guide, not a promise: how far one card goes depends on how often your people actually ask.
Figures based on published model sizes. Last reviewed September 2026.
Your case will not match a table exactly, which is why the scoping call starts here.
Scoped per engagement, quoted before anything is bought. Hardware and service are priced separately, so you can see what each one costs.
The service fee is set by three things: how many people will use it at once, how large a model it has to run, and how much redundancy you need. The scoping call answers all three.
Tell us what you want AI to do and on which data, and we will come back with a sized specification and a quote.