Artificial Intelligence

What Is Ollama and How Do You Install It on Your Own Server?

Talha Aslan 19 min read 3 views

What is Ollama and what does it do?

Ollama is an open source tool that lets you download and run large language models (LLMs) on your own computer or server. It pulls model files, loads them into memory and exposes them through a command line and a local REST API. As a result, your prompts never leave your machine for a third-party cloud.

We are Talha Aslan and team, a digital marketing, web and AI automation team. We are not a hosting company. That is why every command and setting name in this guide comes from the official documentation and the GitHub repository of the project. Our goal is simple: help you decide when a self-hosted model server makes sense, and how to set one up safely.

If the LLM concept is new to you, start with our guide to large language models in software development. Also note that this article covers the model server only. If you want an agent that uses tools on top of a model, our self-hosted AI agent guide covers that layer.

Why run a language model on your own server?

There are three main reasons to run models locally: privacy, cost control and independence from the internet. However, each benefit comes with a trade-off. Knowing both sides up front saves you from the wrong expectations.

  • Privacy: According to the official FAQ, the developers do not see your prompts or data when you run models locally. Customer emails, internal documents and personal data stay on your server.
  • Cost control: Cloud APIs usually charge by token usage. On your own server, cost shifts to fixed items such as hardware and power. For heavy and predictable use, that is easier to budget.
  • Offline use: Once a model sits on disk, it answers without an internet connection. That matters for closed networks and field devices.
  • Customization: You decide the system prompt, parameters such as temperature, and the context length.

On the other hand, the cost of ownership is real. You handle hardware, updates and security. In addition, the open models you can run locally may not match the largest commercial cloud models on every task. In short, this is not a "free ChatGPT". It is an infrastructure component, and the responsibility moves to you.

How much RAM and VRAM does a local LLM need?

We will not give you one fixed number, because the answer depends on the parameter count, the quantization level and the context length. The rule itself is simple, though. Above all, the model weights must fit in memory, and you need extra room for the context on top.

Example calculation (rough estimate, not a guarantee): Take a model with 8 billion parameters at 4-bit quantization. Each parameter needs about half a byte. So the weights alone come to roughly 4 GB. The context window and working buffers come on top of that. The 16-bit version of the same model needs about four times as much.

This math gives you three practical rules:

  • When the parameter count doubles, memory needs roughly double too.
  • Lower-bit quantization saves memory. That said, it can cost some answer quality.
  • A longer context also uses more memory, so plan for it if you summarize long documents.

Specifically, every model page in the library lists the download size. That size is a good hint for the lower bound of memory you need. Therefore, check the target model size before you pick a server, then add headroom.

Can you run a local model without a GPU?

Yes. If the tool cannot find a supported GPU, it runs the model on the CPU and system RAM. However, generation gets noticeably slower. Still, a CPU can be enough for testing small models, overnight batch jobs or low-traffic internal tools.

For real-time chat, many users or large models, a GPU makes the difference. The official Docker documentation gives separate run commands for NVIDIA and AMD cards. Also, if a model does not fit in GPU memory, the server can split it between GPU and CPU. In that case, speed lands somewhere in between.

You can check which processor a model uses with ollama ps. According to the official FAQ, this command lists loaded models and their processor split: 100% GPU, 100% CPU or a mix. Our GPU server rental guide helps at this point.

If you are unsure about the server type, our VPS vs cloud server vs VDS comparison walks through the basic options. Put simply, define the workload first and choose the hardware second.

How do you install Ollama on a Linux server?

The official Linux documentation offers an install script. The docs show it as a one-liner that downloads the script and pipes it straight into the shell. We do not recommend that pattern, because you would run a script from the internet with admin rights without reading it. Instead, follow three steps: download, inspect, run.

curl -fsSL https://ollama.com/install.sh -o install.sh
less install.sh
sh install.sh

In the second step, you then read what the script does. In general, it places the binary, creates an ollama system user and sets up a systemd service. It also asks for admin rights where it needs them.

If you prefer not to use the script, the docs also describe a manual install. You download the archive, extract it under /usr and then write the service file yourself. The unit file in the docs looks like this:

[Unit]
Description=Ollama Service
After=network-online.target

[Service]
ExecStart=/usr/bin/ollama serve
User=ollama
Group=ollama
Restart=always
RestartSec=3

[Install]
WantedBy=multi-user.target

A manual install gives you more control. On the other hand, you also handle updates by hand. So for a single test server, the reviewed script route is usually more practical.

How do you check the service after installation?

The first job after installation is to confirm that the service runs. The order in the docs goes like this: reload the systemd configuration, enable the service at boot, then check its status.

sudo systemctl daemon-reload
sudo systemctl enable ollama
sudo systemctl start ollama
systemctl status ollama

If the service fails, read the logs with journalctl -e -u ollama. For example, messages about a missing GPU driver, low memory or a port in use show up there.

Finally, test the API from the same server. The endpoint that lists local models works well for this:

curl http://localhost:11434/api/tags

An empty list is normal if you have not pulled a model yet. What matters is that the connection works. For updates, the docs suggest running the install script again. Here, too, stick to the download and inspect routine.

To uninstall, the docs give this order: stop the service, disable it, delete the unit file, then remove the ollama user and group. After that, delete the directory that holds the models. Think twice before you delete it. Models can be large, and downloading them again takes time.

How does setup differ on macOS and Windows?

Desktop setup is simpler. In practice, you grab the app for macOS or the installer for Windows from the official download page. After setup, the app runs in the background and you use the same commands in a terminal.

Instead, the differences show up mostly in configuration. According to the official FAQ:

  • macOS: You set environment variables with launchctl setenv and then restart the app. Models live under ~/.ollama/models by default.
  • Windows: You edit user environment variables in the system settings, then reopen the app from the Start menu. Models live in the .ollama\models folder inside your user profile.
  • Linux: You pass settings through a systemd override file. Models live under /usr/share/ollama/.ollama/models by default.

A desktop install is ideal for testing models and designing prompts. However, for a service your whole team shares, use a server instead of keeping your laptop awake. That way, sleep mode, OS updates and personal use do not interrupt the service.

Which command runs Ollama in Docker?

First, the official image comes from the project's own repository on Docker Hub. The CPU-only command in the docs stores models in a named volume and publishes port 11434 with -p 11434:11434. Thanks to the volume, your downloaded models survive when you remove and recreate the container.

On a server with an NVIDIA card, you add the --gpus=all flag. For that, the NVIDIA Container Toolkit has to be in place on the host. For AMD cards, the docs show the rocm image tag plus device flags.

Pay attention to one detail. When you write -p 11434:11434, Docker publishes the port on every network interface of the server. Moreover, Docker's own network rules can bypass firewall tools such as ufw in some setups. So if only apps on the same server use the API, bind the port to localhost:

docker run -d -v ollama:/root/.ollama -p 127.0.0.1:11434:11434 --name ollama ollama/ollama
docker exec -it ollama ollama run gemma4

The second line pulls a model inside the container and starts an interactive session. If Docker is new to you, read our Docker containers guide first. It explains volumes, ports and images in detail.

How do you download and run a model?

Day-to-day work comes down to a handful of commands. In the examples below, we use gemma4, the model name in the official CLI reference. You replace it with the model you pick.

ollama pull gemma4
ollama run gemma4
ollama ls
ollama ps
ollama stop gemma4
ollama rm gemma4

In other words, here is what each command does:

  1. pull downloads a model from the library without starting it. You do this first when you prepare a server.
  2. run starts the model and opens an interactive chat. If the model is missing, it downloads it first.
  3. ls (long form list) shows the models on disk.
  4. ps shows the models that sit in memory right now, plus the GPU and CPU split.
  5. stop unloads a model from memory right away.
  6. rm deletes a model from disk.

In addition, ollama create builds your own model definition from a Modelfile. For example, you can define a "customer support" variant with a fixed system prompt and temperature. Then your app does not need to send the same settings with every request.

Which tasks suit a local model best?

A local model does not replace a cloud model for every job. Still, for some tasks it delivers enough quality and keeps data in-house. In our AI automation projects, we suggest these starting points:

  • Classification: Tag incoming support tickets, form messages or reviews by topic and priority.
  • Summarization: Turn meeting notes, long email threads or draft reports into short bullet points.
  • First drafts: Prepare product descriptions, category copy or reply templates. A human still does the final check.
  • Data extraction: Pull names, dates and amounts from invoices, orders or applications into a structured format.
  • Semantic search: Use embedding models to turn internal documents into vectors and answer "which document is similar to this one?".

What these jobs share is short output that a person can review. On the other hand, long, creative and error-free text may call for a bigger model. Therefore, pick a small and measurable first project. Once it works, widen the scope.

Which model should you choose?

In practice, model choice takes more time than installation. We do not name one "best" model, because the right pick depends on your task, your language and your hardware. Instead, narrow it down in this order:

  1. Task: Chat and text generation, code, image understanding or embeddings for search? The library tags models by these capabilities.
  2. Language: If you work in more than one language, test multilingual quality with your own sample texts.
  3. Size: Start with the largest size that fits comfortably in your server's memory. Then scale down if needed.
  4. License: Read the commercial terms on the model page. We cover this in its own section below.

A practical method works well here. Write 20 realistic questions from your own business and run two or three candidates against the same set. After that, note answer quality, speed and memory use. This small test gives you a sounder decision than any general leaderboard.

For example, an online store that wants product description drafts may do fine with a mid-size chat model. By contrast, contract summaries with long context need more memory and careful testing.

How do you connect the REST API to your app?

As soon as the service runs, it opens a local HTTP API. According to the official docs, the local base URL is http://localhost:11434/api. You use /api/chat for conversations, /api/generate for one-off completions and /api/embed for embedding vectors.

curl http://localhost:11434/api/chat -d '{
  "model": "gemma4",
  "messages": [{"role": "user", "content": "How does our return policy work?"}],
  "stream": false
}'

Specifically, the "stream": false value returns the answer in one piece. Chat interfaces usually keep streaming on, so users see the first words without waiting.

The tool also offers an OpenAI-compatible endpoint at http://localhost:11434/v1. That means you can point an app built with the OpenAI library at your local model by changing the base URL. We cover the cloud side in our OpenAI API guide, which helps when you compare both options.

One warning matters here. The official docs state it plainly: local requests need no authentication. In other words, anyone who can reach the API can use your models. The next section deals with exactly that risk.

Why is exposing the Ollama API to the internet risky?

According to the official FAQ, the server binds to 127.0.0.1 on port 11434 by default. So right after setup, only the same machine can reach it. That is a safe default. The trouble starts when you set OLLAMA_HOST to 0.0.0.0 and open the port to the internet.

Because the local API asks for no password or key, an open port brings these risks:

  • Resource abuse: Strangers use your CPU and GPU for their own jobs. You pay the bill and suffer the slowdown.
  • Model management: The API covers more than chat. It also has endpoints to pull and delete models, so an outsider could fill your disk or remove your models.
  • Configuration leaks: The system prompts and settings of your custom models become readable from outside.

Therefore, the rule is clear: never expose this port without protection. If you need remote access, put a reverse proxy with authentication in front of it, restrict access by IP, or use a VPN. Also run a host firewall. Our CSF firewall guide is a good place to start.

How do you protect the model server with an Nginx reverse proxy?

The most common setup keeps the model server on localhost and puts Nginx in front of it. Nginx talks HTTPS to the outside world, asks for basic authentication and accepts only the IP addresses you allow. The official FAQ shows an Nginx example that forwards requests to the local port and sets the Host header to localhost:11434.

server {
    listen 443 ssl;
    server_name llm.example.com;
    # ssl_certificate and ssl_certificate_key lines go here

    location / {
        allow 203.0.113.10;
        deny all;
        auth_basic "LLM";
        auth_basic_user_file /etc/nginx/.htpasswd;
        proxy_pass http://127.0.0.1:11434;
        proxy_set_header Host localhost:11434;
    }
}

Next, you create the password file with the htpasswd utility. The IP address in the example comes from a documentation range. Replace it with the address of your office or app server.

If you set up Nginx from scratch, follow our Nginx install and server block guide. We explain the reverse proxy concept and a GUI alternative in our Nginx reverse proxy guide. After you install the certificate, check your domain with our SSL checker.

What do web interfaces like Open WebUI add?

The command line and API are enough for developers. However, colleagues outside the dev team prefer a chat screen in the browser. Open WebUI is an open source web interface that can connect to this API. It adds user accounts, chat history and model selection.

Follow the official Open WebUI documentation for its setup, not a third-party tutorial. The project keeps versions, image names and environment variables current on its own pages. That is why we do not list commands for it here.

Adding an interface changes your security picture. You now run two services: the model server and the interface. Keep these principles:

  • Keep the model server on localhost or an internal Docker network. Expose only the interface.
  • Create the first admin account right away. Then turn off open sign-ups or require approval.
  • Serve the interface over HTTPS as well.

This way, users chat in the browser while the model server stays out of direct reach.

How do you manage memory and concurrent requests?

You tune the server through environment variables. On Linux, the official route is a systemd override file:

sudo systemctl edit ollama
# in the editor that opens:
[Service]
Environment="OLLAMA_KEEP_ALIVE=10m"
Environment="OLLAMA_MODELS=/data/llm-models"

sudo systemctl daemon-reload
sudo systemctl restart ollama

These are the main settings from the official FAQ:

VariableWhat it controlsWhen to change it
OLLAMA_HOSTBind address and portWhen the proxy runs on another machine; still never expose it unprotected
OLLAMA_KEEP_ALIVEHow long a model stays in memory (5 minutes by default, per the FAQ)To cut first-response delay or free memory sooner
OLLAMA_MODELSModel storage directoryWhen the system disk is small and you want a separate data disk
OLLAMA_NUM_PARALLELParallel requests per modelFor multi-user access; it raises memory use
OLLAMA_MAX_LOADED_MODELSModels loaded at the same timeWhen you switch between several models often
OLLAMA_CONTEXT_LENGTHDefault context windowFor long documents; it increases memory needs

In short, every setting is a trade-off. Parallel requests and long context improve the user experience, but they eat memory. After each change, watch the result with ollama ps.

Do open model licenses allow commercial use?

Two separate licenses apply here, and people often mix them up. The software itself carries the MIT license in its GitHub repository. However, every model you download has its own license, and that license belongs to the model's developer.

Some models come under permissive open source licenses such as Apache 2.0. Others ship with the developer's own terms of use. Those terms can include an acceptable use policy, extra permission above a certain user count, or limits on using outputs to train other models.

In practice, take these steps:

  1. First, open the license section on the model's library page.
  2. Read the full text on the developer's official site.
  3. If you plan a commercial product, a client-facing service or fine-tuning, note whether the license allows it explicitly.
  4. Where you are unsure, ask your legal counsel.

This section is not legal advice. We simply want to point out that "open model" does not always mean "free for any use". Also, if you process personal data, your privacy obligations under laws such as GDPR still apply, no matter where the model runs.

Security checklist for a self-hosted model server

Once setup is done, tick off this list item by item, because small gaps add up. The points come from the official docs and from general server security practice.

  • Does the model server listen only on 127.0.0.1 or an internal address? If you use Docker, is the published port bound to localhost?
  • Is remote access possible only through an authenticated reverse proxy, a VPN or an IP allowlist?
  • Does the host firewall open only the ports you need?
  • Did you read the install script before you ran it, and did it come from the official domain?
  • Do you update system packages and the model server on a regular schedule?
  • Is there enough disk space for models, and do you monitor it?
  • Do the licenses of your models fit your business use?
  • If you run a chat interface, did you turn off open sign-ups?

We suggest you review this list once a month. Each new interface, plugin or agent widens your attack surface. For broader web application risks, our OWASP Top 10 guide is a useful companion.

Ollama or a cloud LLM API: which fits better?

Above all, both approaches have different strengths. The table below is a general summary to support the decision. It does not reflect the pricing or performance of any specific provider.

CriterionSelf-hosted (your server)Cloud LLM API
Where data goesStays on your serverTravels to the provider's infrastructure
Cost modelHardware and running costs, largely independent of usageUsually pay per use
SetupServer, install and security are on youStart right away with an API key
Model choiceOpen-weight modelsThe provider's own models, including the largest ones
ScalingManual, by adding hardwareHandled by the provider, within quota limits
Offline usePossibleNot possible
Maintenance and updatesYour responsibilityThe provider's responsibility

For many teams, the answer is a mix. You process sensitive data with a local model and send the jobs that need top quality to a cloud API. The OpenAI-compatible local endpoint makes that switch easier.

When should you not set up Ollama yourself?

To be honest, not every business needs its own model server. In these cases, we suggest you wait or leave the job to a hosting provider or a technical team:

  • Nobody on your team can handle Linux updates, firewall rules and log checks.
  • Your usage is a few dozen requests a month. Then the server bill often exceeds what a cloud API would cost.
  • Your work needs the quality of the strongest commercial models, and your tests with open models fell short.
  • You must promise high availability. That calls for an architecture beyond a single server.
  • You have a shared hosting plan. These plans usually do not allow long-running services or high memory use.

Even then, testing the tool on your own computer is worth it. You learn which model fits your work, and then you make the server decision with data.

Where should you start?

Our suggested order goes like this. First, try a small model on your own computer. Next, compare two or three models with real business questions. Then choose a server that fits, run the reviewed install script and keep the API on localhost.

When you need remote access, add Nginx and authentication, check the licenses and review the security list on a schedule. If you follow this order, your local model server becomes a solid AI component that keeps your data in-house.

Want to connect a local model to your workflows, for example to sort support tickets or draft product copy, our AI automation services team plans that integration. Likewise, if you need a custom interface or workflow for your own software, we can also help through custom software development.

Sources: Ollama Linux install docs, Ollama FAQ, Ollama Docker docs, Ollama GitHub repository.

Frequently Asked Questions

Is Ollama free?
Yes, the software is free and open source under the MIT license in its GitHub repository. However, you pay for the hardware, server rent and power that run the models. Also, every model you download has its own license, so read the terms on the model page before any commercial use.
Does Ollama work without an internet connection?
Yes. Once a model sits on disk, it generates answers without an internet connection. You only need the internet to download or update models, or to use optional cloud models. That makes local models attractive for closed networks, field devices and teams that handle sensitive data.
Which port does Ollama use by default?
According to the official FAQ, it binds to 127.0.0.1 on port 11434 by default. So right after setup, only the same machine can reach it. You can change the address with the OLLAMA_HOST variable, but the local API has no authentication, so never expose the port without protection.
Can I run Ollama on shared hosting?
Usually not. Shared hosting plans rarely allow long-running background services, custom ports or high memory use. You need a VPS, a dedicated server or a GPU cloud instance with admin access. If you are unsure, ask your hosting provider about their rules for running services.
Can I reuse my OpenAI API code with Ollama?
Largely, yes. It offers an OpenAI-compatible endpoint at http://localhost:11434/v1. In an app built with the OpenAI library, you change the base URL and model name to point requests at a local model. Still, do not assume every feature works the same; test the endpoints you rely on.
How long does Ollama keep a model in memory?
According to the official FAQ, a model stays in memory for five minutes after its last use by default, then unloads. You can extend or shorten that window with the OLLAMA_KEEP_ALIVE environment variable. To free memory right away, use the stop command with the model name.
  • Ollama
  • Local LLM
  • Large Language Models
  • Self-Hosted AI
  • Docker
  • Nginx
  • Open Source AI
Share:
Talha Aslan

Google Partner digital marketing expert. Hands-on with SEO, Google Ads, web design and e-commerce projects since 2012; every post here comes from that experience.

Next project

Let's talk about your project.

Your brief goes straight to Talha Aslan and team: strategy led by Talha, delivery by an experienced team. The first consultation is free; we listen and come back with a clear roadmap.