Skip to content

What is a self-hosted LLM?

A self-hosted LLM is a large language model that runs on hardware a company controls. The hardware stands in the company’s own server room or at a hosting provider under the company’s contract. Prompts and documents stay on those machines. They do not reach the company that trained the model.

How a self-hosted LLM works

Self-hosting starts with an open-weight model. The maker publishes the trained weights, and a company downloads them. DeepSeek, Qwen and Mistral are model families with published weights.

Serving software such as vLLM or Ollama loads the weights and answers requests through an API. vLLM serves a model in the API format of OpenAI (vLLM documentation). An application that speaks this format moves to the self-hosted model with a new address and a new key.

Open weights are narrower than open source. The company receives the trained model, usually without the data it was trained on. The license differs from maker to maker. The license check comes before the hardware.

Self-hosted LLM or cloud API

QuestionSelf-hosted LLMCloud API
Where does the model run?On servers the company controlsOn the vendor’s servers
Where do prompts and documents go?They stay on the company’s serversTo the vendor
Which models are there to choose from?Models with published weightsThe models of the vendor
Who supplies the hardware?The company buys or rents itThe vendor
Who runs the model?The company’s IT team or its hosting providerThe vendor

A company can use both. Data that must not leave the company goes to the self-hosted model. Other tasks stay with a cloud model.

An open model is good enough for some tasks and falls short in others. Bitkom notes that many tasks do not need the largest and newest model (Bitkom). A test with the company’s own tasks shows where the self-hosted model is enough.

What the operation needs

The hardware depends on the size of the model and on the number of people who ask at the same time. Long documents add to the memory a model needs. An assessment of the tasks comes before the purchase.

The model is one part of the operation. The other part decides who may ask it what:

  • a token per application and a log of the requests,
  • a second machine or an approved cloud model that takes the requests when the first server is down,
  • a plan for updates, so a new model runs beside the old one until the staff approve it,
  • a person who runs the model, with a runbook for restarts and failures.

The model reaches the company’s systems through an API or an MCP server with a short list of tools. On-premise AI in a server room or a colocation rack adds the questions of the room: power, cooling, network and physical access.

What I bring to a self-hosted LLM

I am an IT expert for data platforms. The operation around a model is server work, and I have done that work with my own hands. In data centers in Germany and the United States I set up servers, switches, firewalls and backup systems and did the cabling. I have run blockchain nodes in several data centers. I moved infrastructure from one cloud to another with failover.

In a client project I rebuild a platform with Claude and Codex under one rule file. The rules and the memory of the agents stay outside the model. An open-weight model on a company’s own servers can take a cloud model’s place in that setup. The agents can run on an open-source control plane for AI agents, which a company installs on a server it chooses.

For companies in Germany I plan and set up self-hosted LLMs on servers they control. We sort the data by what may leave the company, check the license of the candidate models and size the hardware for the users. I take on this work as a freelancer or as a permanent employee.

Searches this page answers

  • what is a self hosted llm
  • what is self hosted llms
  • what does self hosted mean
  • self hosted large language model
  • self hosted ai server
  • run llm on own server
  • local llm vs cloud
  • what is a private llm
  • what is an on premise llm
  • what is an open weight model