What air-gapped AI actually requires
On-prem AI is frequently promised and rarely delivered. The practical requirements when data genuinely cannot leave your perimeter.
A different set of constraints
For many institutions in the Kingdom, the question is not whether AI is useful. It is whether a given approach is permissible at all. Where data residency, sovereignty and regulatory obligations apply, an architecture that sends prompts to someone else's infrastructure is not a trade-off to be negotiated — it is simply out of scope.
What on-prem really involves
Running models inside your own perimeter changes the engineering problem rather than removing it.
- Infrastructure — GPU capacity, and a platform to schedule and share it across teams
- Model operations — versioning, evaluation and rollback, without a vendor doing it for you
- Data pipelines — retrieval that respects existing access controls rather than flattening them
- Observability — knowing what was asked, what was returned, and on what basis
- Lifecycle — a plan for the day the model needs replacing
The honest trade-off
Air-gapped deployment costs more effort than calling an API, and the capability ceiling moves more slowly. What it buys is the ability to use AI at all in environments where the alternative is not using it. For a bank weighing a supervisory conversation against a convenience, that is usually the entire argument.
The work is ordinary engineering, done carefully. That is precisely why it is worth doing properly.
