# Self-hosted LLM

Source: https://ordergroup.co/glossary/self-hosted-llm/
Last updated: 2026-10-10

> Self-hosted, on-premise and private LLMs for teams buying AI features: GPU memory, serving with vLLM, open models, GDPR and what the system must handle.

[AI](https://ordergroup.co/glossary/ai/)
4 min read

# Self-hosted LLM

Also known as: on-premise LLM, private LLM, local LLM, self-hosted large language model

Definition

A self-hosted LLM is a large language model whose weights run on infrastructure you control, your own servers or rented GPUs. Prompts and outputs never reach a model vendor.

Cite this entry

Text
"Self-hosted LLM". Order Group, Software glossary, 10 October 2026. https://ordergroup.co/glossary/self-hosted-llm/
HTML
`<a href="https://ordergroup.co/glossary/self-hosted-llm/">Self-hosted LLM</a> - Order Group`

Reviewed by [Łukasz Gajownik](https://ordergroup.co/authors/lukasz-gajownik/), Head of AI & Frontend
Last reviewed 10 October 2026

## How a self-hosted LLM works

A self-hosted LLM has three parts: the model weights, an inference server, and the hardware they run on. The weights are files published by the model's maker under a license, for example Google's Gemma or the Polish Bielik and PLLuM models. The inference server loads the weights into GPU memory, accepts requests and generates the answer token by token. Your application calls that server over HTTP, the same way it would call a vendor's API.

vLLM is a widely used open-source inference server. It stores the attention key-value cache in fixed-size blocks, a method its authors called PagedAttention (SOSP 2023), so one GPU can serve many requests at once without reserving memory for the longest possible answer. It also exposes an HTTP API compatible with the OpenAI specification. Code written for a hosted API can usually be pointed at a self-hosted model by changing the base URL and model name.

Model size and precision decide the hardware. Google's documentation puts the memory needed for Gemma 4 26B A4B at about 57.7 GB in 16-bit, 28.8 GB in 8-bit and 14.4 GB in 4-bit quantization. That is before the cache for active conversations, which grows with the number of parallel sessions and the length of the texts. Lower precision saves memory and can cost quality, and that cost has to be measured on your own task.

Mixture-of-experts (MoE) models split some layers into many expert sub-networks and route each token through a few of them. Gemma 4 26B A4B has 26 billion parameters and activates about 4 billion per token. Generation is faster than in a dense model of the same size, but all 26 billion parameters still have to sit in GPU memory, so the hardware is sized for the full model.

The terms on-premise, private and local LLM are used loosely. On-premise usually means the organization's own data center, private often means a dedicated deployment at a cloud provider, and local can mean a laptop. For a buyer the question is the same in each case: who operates the machine that sees the prompts, and where it stands.

## What a self-hosted LLM means for your software

With self-hosting, your team takes on the operations a model vendor would otherwise run. Requirements for any system that runs its own model:

- GPU capacity is sized from concurrency. The number of parallel sessions and the length of prompts and answers set the memory for the cache; a load test with realistic documents gives the real figure.
- The model version is pinned and changes like a release. A new model, quantization or system prompt runs through the same evaluation set as the old one before users see it (see [LLM evaluation](https://ordergroup.co/glossary/llm-evaluation/)).
- The inference server has a health check and comes back on its own. Rented GPU instances can be reclaimed by the provider, and when nothing starts a new one the AI feature is down.
- Requests have timeouts, a maximum input size and a queue limit, and the user gets a clear message when the model is busy.
- Prompts and outputs are personal data as soon as users paste their own documents. Logging, retention and anonymization are decided before the first log line is written (see [data anonymization](https://ordergroup.co/glossary/data-anonymization/)).
- The license of the weights is checked for commercial use and recorded together with the model version.
- Someone owns updates. The inference server, GPU drivers, base image and model weights each have their own release cycle.

Writing the specification?

Add Self-hosted LLM to your requirements checklist

Collect the terms your project touches and get their system requirements in one e-mail, ready for an RFP.

Hosted API, rented GPUs and on-premise compared
AspectHosted model APISelf-hosted on rented GPUsSelf-hosted on-premise

Who processes promptsVendor, under its data processing termsYour software on the GPU provider's machinesOnly your organizationModel choiceVendor's catalogAny model whose license allows itAny model whose license allows itCostPer tokenPer GPU hour, also when idleHardware purchase plus operationsScalingVendor's rate limitsMore instances, if the provider has themHardware you ownOperationsVendorYour team: server, drivers, restartsYour team, hardware includedModel changesVendor can retire a versionYou decide when to upgradeYou decide when to upgrade

## Rules and standards

The GDPR does not require self-hosting, but it shapes the decision. A provider that processes prompts on your behalf is a processor under Article 28 and needs a contract with the terms that article lists. Article 32 names pseudonymization and encryption among the security measures a controller considers. Chapter V restricts transfers of personal data outside the European Economic Area, so the location of the GPUs, and of any support staff with access to them, matters as much as the contract. Open-weight models come with their own licenses, some with use restrictions, and the license belongs in the procurement review next to the hosting contract.

## From our projects

For PSONI, a Polish association for people with intellectual disabilities, we are building Generator ETR, a web and mobile app that turns complex text into [easy-to-read](https://ordergroup.co/glossary/easy-to-read/) Polish. The project specification calls for a locally hosted model: Gemma 4 26B A4B served by vLLM, sized for up to 45,000 requests a month and 100 parallel sessions. During the project we also tested two Polish models, Bielik and PLLuM 12B.

Our deployment scripts start a GPU instance, install vLLM with the Gemma 4 weights and expose an OpenAI-compatible API. Before PSONI's acceptance tests in September 2026 we found that a reclaimed GPU instance was not replaced automatically, which left the staging environment without a model. We switched to fixed GPU capacity as a temporary measure. As of October 2026, the target production infrastructure for the model is still being analyzed.

## Related terms

- [Data anonymization](https://ordergroup.co/glossary/data-anonymization/)

Data anonymization is processing personal data so that no one can identify the person by any means reasonably likely to be used. Anonymous data falls outside the GDPR; pseudonymized data, which can be re-linked with separately kept information, remains personal data under the GDPR.
- [Easy-to-read (ETR)](https://ordergroup.co/glossary/easy-to-read/)

Easy-to-read
Easy-to-read (ETR) is a way of writing information so that people with intellectual disabilities can understand it, with everyday words, short sentences and one idea per line, checked by readers from that group. Inclusion Europe publishes the European standard.
- [LLM evaluation](https://ordergroup.co/glossary/llm-evaluation/)

LLM evaluation is the systematic testing of a language model's outputs on a fixed set of inputs against agreed acceptance criteria, using automatic checks, model-based grading and human review. It decides whether a model, prompt or version is good enough to ship.

## Sources

1. [Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP 2023)](https://arxiv.org/abs/2309.06180) - Kwon et al., arXiv
2. [vLLM: OpenAI-Compatible Server](https://docs.vllm.ai/en/latest/serving/online_serving/) - vLLM project
3. [Gemma 4 model overview](https://ai.google.dev/gemma/docs/core) - Google
4. [Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer](https://arxiv.org/abs/1701.06538) - Shazeer et al., arXiv
5. [Regulation (EU) 2016/679 (GDPR)](https://eur-lex.europa.eu/eli/reg/2016/679/oj) - EUR-Lex

Łukasz Gajownik reviewed this entry. Ask how it applies to your project.

[Ask an engineer](https://ordergroup.co/contact-us/)

## FAQ

![Łukasz Gajownik](https://ordergroup.co/media/images/T02DHCC1Z-UT420UBA7-d74022d27c3c-512.format-webp.webp)

Łukasz Gajownik

Head of AI & Frontend

[Talk to an engineer](https://ordergroup.co/contact-us/)

### Does a self-hosted LLM make us GDPR compliant?

No. It removes the model vendor from the list of parties that see the data, but the legal basis, retention, access control and logs are still your responsibility. If the GPUs are rented, the GPU provider is still a processor.

### How much GPU memory do we need?

Start from the model size and precision: about 2 bytes per parameter at 16-bit, half that at 8-bit. Then add the cache for your expected number of parallel sessions and confirm the figure with a load test on real documents.

### Are open models good enough compared with hosted ones?

For a narrow task a smaller open model is often good enough, and for some tasks it is not. Only an evaluation on your own test set answers the question for your system.

### Can we move from a hosted API to a self-hosted model later?

Yes, if the application talks to the model through an OpenAI-compatible API and you keep an evaluation set. The prompts usually need tuning for the new model, and the evaluation shows when they are ready.

Building a system that depends on Self-hosted LLM?

See how we build software for this domain, with case studies and the stack we use.

[See AI Software Development Services](https://ordergroup.co/ai-software-development/)

Requirements checklist

For each term we send the definition and what it requires from your software. Free, no sales call needed.
Your checklist is empty. Use the plus next to a term to add it.
