Most private LLM deployment projects start with one download command and look impressive in week one. Trouble arrives later: when a second department connects, when the first model update lands, or when an auditor asks who saw which documents. This guide covers the decisions that belong before any server is ordered: which work, which model, which hardware and which access rules.
Technical terms are explained briefly where they first appear. The aim is to help you judge whether a quoting team asks the right questions, and whether the setup will still be manageable a year from now.
01Four questions that settle the decision
A private model makes sense when the data cannot leave, the workload repeats every day and someone will own the server. If one of those three is missing, look at the cloud route first. You can settle the matter in a single meeting by writing down answers to the questions below.
- Data class: Do the documents contain personal data, trade secrets or client confidentiality clauses; if so, where do they live today and who can open them.
- Frequency: Is the task daily, or a few times a month; a dedicated server for occasional work sits idle most of the time.
- Ownership: Will an IT person watch patches, disk space and alerts, or will an outside team do it under a maintenance contract.
- Quality bar: Does the result depend on the strongest model on the market, or is it work where a model that quotes, classifies or drafts short summaries is enough.
If the data is not sensitive but you want AI connected to your systems, AI integration with cloud models goes live faster. A hybrid design that uses both routes is also possible; the architecture section covers it.
If the answers are vague, it is too early. A few weeks of teams noting which document they wanted help with clears that up and seeds the test set.
02What a local model takes on, by type of company
A local model pays off on bounded, document grounded work; on open creative tasks the gap to cloud models becomes visible. The examples below are not client cases, they are the kinds of work these deployments typically target.
- Law firm: Comparing clauses across contract drafts and finding the relevant passage in past matters, especially where privilege makes cloud processing a debated question.
- Accounting practice: Suggesting ledger codes for bank transaction descriptions and sorting client correspondence by topic.
- Manufacturing plant: Querying maintenance logs for failure history and recurring defects; on a plant network without internet access this is often the only option.
- Healthcare provider: Drafting administrative letters from clinical notes, aimed at paperwork rather than clinical decisions.
- Software team: Explaining code and drafting tests while source code stays in house.
When the real job is extracting fields from scanned invoices, contracts or application forms, the local model becomes the engine of a document processing AI workflow. Text recognition and validation rules then sit before and after the model as separate steps.
For each use case, also write down what the model will never do, such as give legal opinions or make diagnoses; that list feeds the trap questions in the test set.
03The stack layer by layer: serving, gateway, interface
A solid setup has four layers that can be replaced independently; when the model changes, users should notice only through answer quality.
- Serving layer: The software that loads the model into memory and answers requests. Ollama installs easily and suits small teams; vLLM batches concurrent requests and uses GPU memory efficiently on servers with many users; llama.cpp is a lightweight option that can run on machines without a graphics card.
- Gateway: A middle layer that checks identity, picks the model for each request and writes the log; applications talk to it, never to the model directly.
- Document index: A vector database holding document chunks as numeric representations, searchable by meaning rather than exact words.
- Interface: A self hosted chat screen for staff, or a button added to a business application you already use.
Because Ollama and vLLM expose an OpenAI compatible API, many existing tools connect to the local model by changing an address. In a hybrid design the gateway sends requests containing personal data to the local model and routes non sensitive work that needs stronger reasoning to a cloud model. If you need custom screens or an approval panel around it, plan that part as custom software development.
04Choosing a model: size, license and language fit
The right model is not the leaderboard favorite; it is the smallest model that scores well enough on your own test set. Smaller means less memory, faster answers and easier upkeep, so selection runs from small to large, not the other way around.
- Parameter count: A measure of model size. Quality usually rises with size, and so do memory needs and response time.
- License: Open weight models do not all grant the same freedom. Qwen3 is released under Apache 2.0, while Llama 3.1 comes with its own community license and acceptable use policy, including attribution terms.
- Language fit: Public benchmarks lean on English. For documents in other languages or in legal house style, only your own samples show how a model copes.
- Context window: How much text the model reads in one pass; it matters for long contracts, and every extra page claims memory.
- Thinking mode: Some families, Qwen3 for example, can switch on written reasoning before the answer; quality may rise, and so does waiting time.
Build the test set from real questions per use case, expected answers and a few trap questions the model should decline. Run candidates on the same hardware and settings, and score answers with someone who does the work daily.
A second model sits beside the chat model: the embedding model, which makes documents searchable by meaning. If it handles your language poorly, the right passage is never found. A private LLM deployment proposal should name both models, their licenses and test results.
05Rough math for sizing hardware
Three things decide memory: model weights, the context cache and concurrent requests. Rough math points the way; measurement on the target hardware makes the call.
- Weights: At 16 bit precision each parameter takes about 2 bytes, so a model with 8 billion parameters needs around 16 GB for its weights alone.
- Quantization: Storing numbers with fewer bits. Quantizing to 4 bit cuts weight memory to roughly a quarter, at the cost of some quality; your test set shows whether that loss matters for your work.
- Context cache: The memory where the model keeps the text it has read. It grows with document length and with concurrent requests, and on a busy server handling long files it can outgrow the weights.
- Speed metrics: Time to first token shows how long a user stares at a blank screen; tokens per second shows how fast the answer flows.
There are three typical options: a capable workstation for one team, a GPU server on your premises, or a dedicated server rented in your name from a data center. On premises, electricity, cooling and an uninterruptible power supply join the budget; you can estimate the yearly consumption of an always on server with our electricity cost calculator.
With a rented dedicated server, failures and physical security move to the provider, whose access to the machine must be limited by contract. Either way, a chassis with room for a second card makes growth easier.
06Connecting data sources without breaking permissions
Once the model reads company documents, the main risk is permissions: no user should learn, through the model, the contents of a file they cannot open themselves. Connecting sources is therefore not copying files; it is a mapping job that carries access rights along.
The method behind document grounded answers is called RAG, retrieval augmented generation: when a question arrives, relevant chunks are found in the index and handed to the model as sources. Permission filtering must happen during that search; filtering after the answer is written means the content has already reached the model. We describe a setup that answers from company documents with citations on our enterprise knowledge assistant page.
Before connecting each source, put the following in writing:
- Source and owner: File server, document management system, ERP, mail archive or internal wiki, and the person responsible for each one's content.
- Permission mapping: How folder and group rights in the source carry over to the index, and how often they resync.
- Deletion behavior: How quickly a document deleted or archived in the source drops out of the index.
- Exclusions: Folders that never enter the index, such as HR files or medical certificates.
- Current version rule: Which version the model treats as the source when old and new copies of a document coexist.
07Hardening the server during setup
A server in your own building is not safer than a cloud account by default; the decisions made during setup decide that. The points below belong on your acceptance checklist.
- Verified model files: Models come only from the publisher's official repository, and each file's checksum is compared with the published value. For smaller files you can run the check yourself with a hash generator.
- Safe file format: safetensors is preferred; pickle based files can execute code when loaded, so they never come from untrusted sources.
- Identity in front of the API: Ollama listens only to the local machine by default and does not authenticate users. If it must be reachable on the network, an authenticating reverse proxy with encrypted connections goes in front.
- Outbound traffic: Closed at the firewall, opened only in approved update windows to specific addresses.
- Admin access: Admin accounts are personal, protected with two factor authentication and logged separately.
- Instructions hidden in documents: Text inside a file such as "ignore previous instructions" can steer a model, a risk known as prompt injection. That is why the model gets no rights to send email or delete records.
Before handover, ask for three checks in the acceptance record: the API is unreachable from outside, users cannot query another department's documents, and the server cannot reach addresses off the allow list. Skipping them leaves the deployment unfinished.
08Human approval, logging policy and rollback
Reading and drafting can run freely; any action that leaves the company or changes a permanent record should pass a human. Keep that split in a written permission table and update it whenever a new use case is added.
Separate two levels of logging. Metadata is logged for every request: who, when, which model version, how long and how many tokens. Full text logging is kept only for debugging, for a short period and with narrow access, because the prompt and the answer themselves may hold personal data or trade secrets.
These version rules work well in practice:
- Pinning: The model file in production is recorded by its checksum, so a different file cannot be swapped in silently under the same name.
- Versioned instructions: The system prompt, the task description the model sees with every request, is versioned like code and stored with change notes.
- Test set first: A new model or new instructions go through the test set; if the score drops below the previous version, it does not go live.
- Fast way back: The previous model file stays on disk, so rolling back is one configuration change, not a reinstall.
09GDPR, server location and the AI Act
Keeping data in your network does not remove data protection duties. Article 32 of the GDPR requires appropriate technical and organizational measures, and the permission, logging and network decisions in this guide are their concrete form.
- Rented server: The hosting company acts as a processor and needs a contract under Article 28 that limits what it may do, including physical and remote access, disk disposal and incident notice.
- Server outside the EU or EEA: This can count as a transfer under Chapter V, which needs an adequacy decision or safeguards such as standard contractual clauses. The UK GDPR contains equivalent rules for organizations in the United Kingdom.
- Customer facing use: If the model talks directly to people, Article 50 of the EU AI Act requires that they are told they are interacting with an AI system.
- Log retention: Request logs can be personal data, so their retention period follows your existing retention schedule rather than a default setting.
Outside Europe, sector rules on health or financial data may apply similarly; your legal adviser makes the final assessment.
10A six step path from pilot to company wide use
A private LLM deployment should start with one team and a limited document set, and move to each new stage only against a written exit criterion. The order below is designed to defer hardware spending until measurements are in.
- Use cases and test set: Pick the pilot team, define no more than three tasks and build the test set from real examples of those tasks.
- Measure on rented hardware: Run candidate models through the test set on a GPU server rented by the hour; report quality and speed together.
- Hardware decision: Based on the measurements, choose a workstation, an on premises server or a rented dedicated server, and order it.
- Hardened installation: Set up the serving layer, gateway, sign in and logging; attach the security checklist to the acceptance record.
- Pilot use: The team uses it on real work for a few weeks; wrong or incomplete answers are flagged with one click and added to the test set.
- Staged rollout: If the criterion is met, new departments open one at a time, each with its own permission mapping and a short training session.
Write the exit criterion before the pilot starts: a target test set score, an acceptable wait at peak hours and regular use by the team. Criteria written afterwards bend to the result.
11Measuring value: quality, speed and adoption
The value of a private LLM deployment comes down to three questions: are the answers right, are they fast enough and do people actually use it. If one of the three is weak, the other two will not save the project.
- Faithfulness to sources: Whether an answer matches the document it cites, scored regularly on the test set and on random samples from live use.
- Unsourced answer rate: In tasks that must rest on documents, answers given without a source are counted separately; a rising rate points to a problem in the index or the instructions.
- Peak hour performance: Track waiting time at the busiest hours, not the average; users judge the system by its slowest moments.
- Usage by department: See which teams use it how often each week; low use usually means the model was never built into the workflow, not that the model is weak.
- Time per task: Time a few real tasks before the pilot and again at its end.
On cost, recalculate regularly what the same volume would cost through a cloud API; as pricing and open model quality change, the balance can shift.
12Limits and realistic risks of local models
A local model gives control over data, but it has trade offs in capability and operations; name them at the start.
- Made up answers: Like every language model, a local one can state wrong things with confidence. Instructions that force citations, a rule to say "I do not know" when no document supports an answer, and human review on critical work reduce this risk; they do not remove it.
- Complex reasoning: On multi step calculation, long planning or reading many documents together, open models can trail the strongest cloud models.
- Maintenance load: GPU drivers, serving software and the operating system must stay compatible; one mismatched driver update can stop the server.
- Single point of failure: Decide in advance what happens when the only server fails: wait, switch to a spare machine, or route temporarily to an approved cloud model.
- Dependence on one person: If only one person understands the setup, the biggest risk is organizational, not technical. A runbook and handover document are part of delivery for that reason.
None of these is a reason to stop, but each needs an owner and a written plan. Budget a private LLM deployment as living infrastructure, like software licenses, not as a one time project.
13Common private LLM deployment mistakes
Most mistakes come from doing things in the wrong order. Use the six below as a checklist when reviewing proposals.
- Buying hardware first: A graphics card is bought before any measurement, and the model is then squeezed to fit it. Better: run the test set on rented hardware and choose equipment from the results.
- Trusting public rankings: A model that leads English benchmarks can stumble on your documents and house style. Better: compare candidates on a test set built from your own files.
- Exposing a bare API: The serving address is opened to the network without authentication. Better: put an authenticating gateway and encrypted connections in front.
- Indexing every folder: The whole file server goes into the index with no permission mapping. Better: map rights source by source and keep a written exclusion list.
- Keeping full text logs forever: Every prompt and answer is stored for years. Better: separate metadata logs from short lived, access restricted full text logs.
- Updating straight to production: A new model version is installed untested. Better: compare test set scores and keep the previous version on disk for rollback.
14Questions for your implementation team and next step
A good proposal explains the measurement method before naming a model, and the use cases before naming hardware. Ask the teams you speak with the following, and request written answers:
- Which test set and which environment does the hardware recommendation rest on?
- Whose name will the server, accounts, configuration files and model files be in?
- Which runbooks will be handed over so our own staff can run the setup without you?
- What is the maintenance schedule and rollback plan for model updates, driver updates and security patches?
- If we later add a cloud model for some tasks, what changes in our applications?
We run this work as part of our AI automation services: use case list and test set first, then measurement, then the hardware decision and a hardened installation. You pay for hardware or server rental directly to the supplier, and every account and file is handed over to your company.
Starter options by scope are listed under AI automation pricing. Send us three real tasks the model should take on and the document types involved through our contact form, and we will suggest suitable model families, hardware options and a written quote.