Consulting · AI · Delivery

On-Premise AI, Ready From Day One

You want AI on the documents and records that matter, but that data cannot go to a cloud service. The answer is an AI platform that runs on your own hardware, inside your own network, so the data never has to leave. We size it, specify it, install it and hand it over ready to use, then train your team to run it themselves.

Day oneRUNNING AT HANDOVER
Multi-vendorSPECS MORE THAN ONE BRAND CAN MEET
Your networkSENSITIVE DATA STAYS INSIDE
Dev to productionA PATH TO SERVER-CLASS

The question it answers

Could your staff use AI on confidential documents and personal data tomorrow, without any of it leaving your network?

For most organizations the answer is no. Cloud AI services are ruled out for sensitive material, so the AI project stalls, or staff paste confidential text into public chatbots anyway. Buying a GPU machine does not fix that by itself. Someone still has to size it, build a platform that stays up, install models that work in Thai, and hand over something your own team can run.

This is not an argument for running everything locally. Public and low-sensitivity work can still go to the best model in the cloud. This platform is where the work that must stay inside runs. How that split can be decided per request, by the data's own classification, is set out in Your Best Model Is in the Cloud.

This is the AI domain's delivery engagement: the step that turns ‘we cannot use AI on this data’ into a working platform inside your walls.

What you receive

Six deliverables. The hardware is yours, the platform is yours, and so is every password to it.

1 · Sizing

How large a model you need, how much GPU memory it takes, and how many people can ask questions at the same time. Plus the CPU, memory and storage for everything that runs around the model. Written down, with the working shown.

2 · Procurement-ready specification

A specification more than one vendor can meet, written the way public procurement needs it. Buy through your own process or through us. The specification does not tie you to a brand.

3 · A platform built for 24/7

A virtual machine platform on ZFS storage: mirrored system disks, redundant VM storage, snapshots, and backups kept both on the machine and off it, following the 3-2-1 rule. The GPU is passed through to the virtual machines that need it.

4 · AI ready on day one

Model serving, models chosen to work in Thai, a chat interface for staff, a starter for asking questions of your own documents, and an API your developers can build on.

5 · A path from dev to production

Separate environments from the start. When production moves to server-class hardware, the workstation becomes your development machine instead of something you replace.

6 · Handover

A runbook written for your team, hands-on training, and support terms defined in writing: what is covered, when, and how quickly we respond.

How it runs

1
Scoping call. What people will use it for, how many will use it at once, and which data must never leave. The sizing starts from those three answers.
2
Sizing and specification. We turn the answers into a sized, multi-vendor specification your procurement team can use as it stands.
3
Procurement. Through your process or through us. As a Lenovo 360 Authorized partner we supply Lenovo devices and infrastructure solutions, and we can help you coordinate warranty support with Lenovo. The specification itself stays open to any brand that meets it. Hardware lead times are set by the market, not by us, so we tell you the current figure before you order.
4
Build and acceptance. Platform, storage, backups and the AI stack are installed, then tested against the uses you named in the scoping call.
5
Handover and training. The runbook, hands-on sessions for the people who will run it, and the start of the support terms.

What is and is not included

In scope

  • Sizing, specification and procurement support
  • Virtual machine platform, ZFS storage, snapshots and 3-2-1 backup
  • GPU passthrough to virtual machines
  • Model serving, Thai-capable models, a chat interface and an API
  • A starter for asking questions of your own documents
  • Separate development and production environments
  • Runbook, training and written support terms

Not in scope

  • Organization-critical 24/7 production on a workstation (see below)
  • Building your own AI applications, which is a separate AI Development engagement
  • Migrating or cleaning your existing data
  • Integration with internal systems beyond the API
  • Routing requests between cloud and on-premise models by data classification, which is Knowledge Engineering

A workstation is the right start. It is not the right end for everything.

A workstation-class machine is the right answer for development, and for pilot production where a short outage is acceptable. It is not the right answer for AI your organization cannot run without, around the clock. That belongs on server-class hardware, with redundant power and components.

We tell you which one you need before you buy, and the platform we build moves to the server when you get there.

Where your data goes, stated plainly

  • Nothing leaves your network by default. Models, documents and conversations stay on the machine. Connecting it to anything outside is a decision you make explicitly, not a setting someone left on.
  • The credentials are yours. Every administrator password is handed over and documented at handover.
  • Remote support is your choice. If your support terms include it, you enable it, and you can switch it off at any time.

Where it leads

The platform is built to grow. Each of these is a defined next step, and you are free to take any of them yourself or with anyone else.

More users than it was sized for
Re-size and add GPU capacity, or move production to a server. The platform moves with it.
AI the organization now depends on
Server-class production with redundant components, and the workstation kept as the development machine.
Some work could safely use a frontier model
Route each request by the data's own classification: public material to the cloud, confidential material here. How that works.
Answers from documents and databases together
Knowledge Engineering: governed retrieval over your documents and your databases, running on this platform.
Your own AI applications
AI Development: build on the API the platform already exposes.

How we size it

Three questions decide the machine: how large a model the work needs, how much GPU memory that model takes, and how many people will ask questions at the same time. The model is loaded once. The memory left over is what serves the people asking, and each person asking needs their own share of it.

Worked examples

How many people can ask at once, by GPU memory on one card (planning rule: 2 GB per person asking)

  • ~12B model, ~7 GB
  • ~70B model, ~43 GB
  • ~120B model, ~65 GB

People asking at the same moment (not total staff)

Pilot One department Several departments Organization-wide 0 10 20 30 40 0 24 48 96 ~12B~7 GB ~70B~43 GB ~120B~65 GB

GPU memory on one card (GB)

The same figures, as a table

Model size (4‑bit)Model takes24 GB48 GB96 GB
~12B~7 GB71739
~70B~43 GB21
~120B~65 GB10

Assumes 4-bit models, about 2 GB of memory for each person asking, and 10% of the card kept free for the software itself. A dash means the model does not fit, or leaves no room for anyone to ask.

2 GB is a planning figure, deliberately on the safe side. The real amount depends on the model's design and how long each conversation is: from under 1 GB to about 3 GB per person for a typical conversation, and more for long documents. Memory is also only the ceiling. How fast answers come back with many people asking at once is the other limit, and we check both during sizing.

Automated jobs count too: every bot or scheduled task that asks the AI takes a place, the same as a person.

GPU memory decides how big a model you can run and how many can ask at once. CPU and RAM decide how much application you can build around it.

The scope bands on the chart are a rough guide, not a promise: how far one card goes depends on how often your people actually ask.

Figures based on published model sizes. Last reviewed September 2026.

Your case will not match a table exactly, which is why the scoping call starts here.

How it is priced

Scoped per engagement, quoted before anything is bought. Hardware and service are priced separately, so you can see what each one costs.

The service fee is set by three things: how many people will use it at once, how large a model it has to run, and how much redundancy you need. The scoping call answers all three.

Tell us what you want AI to do and on which data, and we will come back with a sized specification and a quote.

Consulting · AI · Delivery

Tell Us Which Data Must Never Leave.

Sized, specified, built and handed over running. The hardware is yours, the passwords are yours, and your data stays where it is.

Book a Scoping Call →

contact@infozense.com  |  +66-82-242-4008  |  Bangkok, Thailand

Lenovo is a trademark or registered trademark of Lenovo.