Self-Hosted AI Agent: How Do You Set One Up on Your Own Server?

What is a self-hosted AI agent, and how do you set one up?
A self-hosted AI agent is software that connects a language model to tools, memory and workflows, running on a server you control. In practice, you install Docker on a VPS, start an orchestrator such as n8n and optionally a local model server such as Ollama, then lock it down with HTTPS, secrets and permission limits.
We are Talha Aslan and team, a digital marketing and web team, not a hosting company. That is why every command and setting in this guide comes from the official documentation of n8n, Ollama, Docker and OWASP. Our goal is simple. We want a developer who runs their own VPS, or a technically minded site owner, to understand each step of the setup.
We covered where agents help in marketing, and which use cases save time, in a separate article: AI agents for marketing and how to build them. This guide does not repeat those use cases. Instead, it focuses on the server side: architecture, installation, security and maintenance.
How does an AI agent differ from a chatbot or classic automation?
A chatbot answers a question with text and usually stops there. Classic automation follows a fixed path. A form arrives, a row goes into a spreadsheet, an email goes out. You define every step, and the system makes no decisions.
An AI agent sits between the two. It receives a goal, uses a language model to decide which tool to call and in what order, then checks the result and picks the next step. For example, it can read a support ticket, query the order system and then draft a reply.
That ability to decide is also the source of risk. The agent makes tool calls that you did not approve one by one. As a result, running an agent on your own server means more than hosting an app. You are operating software whose decisions you need to limit. If you want a refresher on how language models work, see our guide to large language models.
Why would you run an agent on your own server?
The idea of self-hosting usually comes from three needs. A hosted agent platform gets you started fast. However, you have limited control over where data travels, how much you pay and what happens when the platform changes its terms.
- Data privacy: Your workflows, customer records and credentials stay in a database on your own server. With a local model, the prompt text never leaves the machine either.
- Cost control: You pay a fixed server fee. You do not depend on per workflow or per execution pricing, and you track token usage yourself if you call a model API.
- Less lock-in: With an open source orchestrator, you can export workflows, swap the model provider and move the server to another host.
That said, none of these benefits arrives on its own. Data privacy, for instance, only holds while the server stays secure. An exposed admin panel creates a far larger risk than any hosted platform would.
When is self-hosting the wrong choice?
To be honest, self-hosting an agent is unnecessary overhead for many businesses. Even the official n8n installation docs recommend self-hosting for expert users. They also state plainly that mistakes can lead to data loss, security issues and downtime.
We suggest a managed service in these situations:
- Maintenance load: If nobody will track OS updates, Docker images and backups, the server slowly turns into a security hole.
- Model quality: Local models that fit on a small VPS often fall short of large cloud models in reasoning. On complex tasks, the agent makes more mistakes.
- Security ownership: Every incident on your server is your responsibility. That includes intrusion detection, log review and incident response.
- Critical processes: Payments, invoicing and health data leave little room for error. Here you need an experienced team.
In short, your own server is only an advantage if you have the time and skills to run it.
What are the building blocks of a self-hosted AI agent?
It helps to think of a self-hosted AI agent as four layers. Each layer can run in its own container or as an external service. So when you replace one part, the others keep working.
- Orchestrator: This layer runs the agent logic. You can build visually with an open source workflow tool such as n8n, or write code with a Python or JavaScript agent framework.
- Model: This is the language model that makes decisions. You can connect to a cloud provider API or run a local model server such as Ollama.
- Tools and integrations: These are the systems the agent touches. Your CRM, email, calendar, spreadsheets, internal APIs and web search all live here.
- Memory and database: This layer stores workflow definitions, execution history, credentials and, if needed, conversation memory. By default n8n uses SQLite, and it also supports PostgreSQL.
In this guide we use n8n and Ollama as examples, because both are open source and both have detailed official Docker docs. Still, the same architecture applies to a code based framework. What matters is that you limit access for each layer separately.
Cloud API or local model: which one should you pick?
The model layer has the biggest impact on hardware needs and privacy. The table below compares both approaches at a high level. We do not list numbers, because prices and hardware needs depend on the model you choose.
| Criterion | Cloud model API | Local model (such as Ollama) |
|---|---|---|
| Where does processing happen? | Prompt text goes to the provider | Prompt text stays on your server |
| Hardware needs | A small VPS is often enough | Plenty of RAM for the model size, ideally a GPU |
| Model quality | Access to large, current models | Limited to models that fit on the server |
| Cost structure | Usage based token fees | Fixed server and hardware cost |
| Maintenance | The provider maintains the model | You download, update and monitor models |
| Speed | Depends on network latency | Depends on your hardware |
A hybrid setup makes sense for many projects. For example, you summarize sensitive text with a local model and call a cloud model only for steps that need stronger reasoning. In that case, document exactly which data leaves the server.
How much hardware does a self-hosted AI agent need?
The hardware a self-hosted AI agent needs depends on where the model runs. If you only call a cloud API, the server carries the orchestrator, the database and the reverse proxy. A small VPS then covers most workflows, because the heavy computing happens on the provider side.
Running a local model changes the picture. The model loads into memory while it runs. So you need free RAM or GPU memory roughly in line with the model size. The official Ollama Docker docs also ask you to install the NVIDIA Container Toolkit first if you want GPU support. It works without a GPU too, but responses get noticeably slower.
We do not give hard numbers, because requirements vary by model. Instead, follow this order. First, check the size listed on the model page. Next, check the free memory on your server. Finally, run a load test with a real workflow.
We explain the difference between VPS, VDS and cloud servers in our VPS vs cloud server comparison. On small servers that run close to their memory limit, a swap file also helps. Still, swap never replaces real RAM for a local model.
What should you prepare before installation?
Getting the basics ready before you run any command saves you rework later. We recommend you complete this list in order:
- Server: A VPS with a current Linux distribution, accessed by a sudo user rather than root.
- Docker and Docker Compose: Install them from the official Docker package repository. If containers are new to you, read our Docker guide first.
- Domain and DNS: Create a subdomain for the agent and point its A record at the server IP. You can confirm propagation with our DNS lookup tool.
- SSH hardening: Turn off password login and use keys. The steps are in our swap file and SSH hardening guide.
- Model access: If you plan to use a cloud model, create an API key with the provider and set a spending limit.
If you prefer not to work with raw Docker commands, an open source PaaS panel is another option. We covered that route in our Coolify guide. Even so, the security principles below still apply when you use a panel.
What does a Docker Compose skeleton for a self-hosted AI agent look like?
The file below runs n8n and Ollama in one Compose project. We took the image names, port and data paths from the official n8n and Ollama Docker docs. We used example.com in place of a real domain, so add your own values.
services:
n8n:
image: n8nio/n8n
restart: always
ports:
- "127.0.0.1:5678:5678"
environment:
- N8N_HOST=agent.example.com
- N8N_PROTOCOL=https
- N8N_WEBHOOK_URL=https://agent.example.com/
- N8N_PROXY_HOPS=1
- N8N_ENCRYPTION_KEY=${N8N_ENCRYPTION_KEY}
- N8N_ENFORCE_SETTINGS_FILE_PERMISSIONS=true
- GENERIC_TIMEZONE=Europe/London
- TZ=Europe/London
volumes:
- n8n_data:/home/node/.n8n
ollama:
image: ollama/ollama
restart: always
volumes:
- ollama:/root/.ollama
volumes:
n8n_data:
ollama:
Note that we bound the n8n port to 127.0.0.1 only. The official n8n Compose example takes the same approach. For Ollama, we published no port at all. That works because n8n reaches Ollama by service name on the shared Compose network.
In production, pin images to a fixed version tag so you decide when to upgrade. If you want PostgreSQL, start from the PostgreSQL example in the official n8n hosting repository. Once you save the file, run docker compose up -d to start it and docker compose logs -f n8n to follow the logs.
How do you keep secrets in environment variables?
Never write API keys, encryption keys or database passwords straight into the Compose file. Instead, create a .env file in the same folder and reference each value by variable name. That way, your secrets stay private even if you share the Compose file in a Git repository.
# .env (example only, never commit real values)
N8N_ENCRYPTION_KEY=put-a-long-random-value-here
# generate a random value
openssl rand -hex 32
# tighten file permissions
chmod 600 .env
According to the official n8n docs, n8n creates a random encryption key on first launch and saves it in its data folder. It then uses that key to encrypt credentials before writing them to the database. So if you lose the key, you cannot read the credentials in a restored database.
Follow three rules. First, add .env to .gitignore. Second, store the encryption key separately in a password manager. Third, enter cloud model keys through the n8n credentials screen. Also, create a separate key per project with your model provider. If one leaks, you revoke only that key.
How do you add a reverse proxy and HTTPS?
Rather than exposing n8n directly, you put a reverse proxy in front of it. The proxy handles HTTPS on port 443 and forwards traffic to the n8n container, which listens on the local address only. In this setup the official n8n docs ask you to set N8N_PROXY_HOPS to 1 and to forward the X-Forwarded headers.
server {
listen 443 ssl;
server_name agent.example.com;
location / {
proxy_pass http://127.0.0.1:5678;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Host $host;
proxy_set_header X-Forwarded-Proto $scheme;
}
}
The Certbot nginx plugin can add the certificate lines for you with sudo certbot --nginx -d agent.example.com. The Upgrade headers keep the live editor connection open. To install nginx from scratch, see our nginx installation guide. To manage it through a web UI, read our nginx reverse proxy and Nginx Proxy Manager guide. After setup, verify the certificate with our SSL checker.
How do you connect a local model through Ollama?
Once the Ollama container is up, you need to download a model. You pick the model name from the Ollama model library. The commands below use a placeholder.
docker compose exec ollama ollama pull MODEL_NAME
docker compose exec ollama ollama list
Next, open the Ollama credential in n8n and enter http://ollama:11434 as the base URL. That address only works inside the Compose network, so nobody can reach it from outside. The official Ollama FAQ also notes that the server binds to 127.0.0.1 by default. To expose it on a network, you would have to change the OLLAMA_HOST variable.
Our advice is to keep Ollama off the internet entirely. If another machine needs access, put an authenticating reverse proxy in front of it or connect over a VPN. Also check the official docs for the options that control how long a model stays in memory. That way, an idle model does not hold on to RAM it does not need.
Finally, start your first workflow with a small test. For instance, build a flow that only summarizes a piece of text and touches no tools. Once the model responds, add tools one at a time.
How do you protect SSH, the firewall and Docker ports?
The foundation of server security is to keep only the ports you need open. An agent server usually needs SSH, HTTP and HTTPS. On Ubuntu, you can do this with ufw:
sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enable
sudo ufw status verbose
There is an important trap here. The official Docker docs state that traffic to published container ports gets diverted before it reaches the ufw rules. In other words, if you publish a port on 0.0.0.0 in your Compose file, outsiders may reach it even though ufw shows it as closed. That is why our skeleton binds n8n to 127.0.0.1 and publishes no Ollama port.
We also recommend Fail2ban against brute force attempts on SSH. We walk through it in our Fail2ban setup guide. After setup, run a port scan from outside and confirm that only the ports you expect are open.
Why is prompt injection the biggest risk for your agent?
Prompt injection happens when input to the model changes its behavior in unintended ways. The OWASP page on LLM01 Prompt Injection puts this risk first on its list for LLM applications. It also separates two types.
In direct injection, a user steers the model through the text they type into the chat. In indirect injection, the attack hides inside a web page, email or document that the agent reads. While processing that content, the agent may treat the hidden instructions as part of its own task.
This is where an agent differs from a chatbot. A tricked chatbot gives a wrong answer. A tricked agent with tools can send emails, delete records or export data. Moreover, no known method blocks indirect injection completely. OWASP itself frames its advice as layers that reduce risk.
In short, never assume the model will always follow your instructions. Base your security on server and workflow rules that limit what the agent can do, not on the good behavior of the model.
How do you apply least privilege and human approval?
Least privilege and human approval for high risk actions stand out among the OWASP recommendations. The same list also names excessive agency, meaning an agent with more functions than its task requires, as a separate risk. In practice, you can take these steps:
- Create a separate read only account or API key for each integration. Grant write access only to the workflow that truly needs it.
- Add an approval step before irreversible actions such as deletes, payments, bulk email and external sharing.
- Narrow the file folder the agent can reach. The official n8n Compose example uses a variable that limits file access to a single folder.
- Keep outside content, such as web pages, emails and documents, apart from system instructions and label it clearly.
- Check that the agent output matches the expected format before you pass it to the next tool.
These rules slow the agent down, and that is fine. In the first weeks, you can review the approval logs to see which actions are safe to automate. Then you widen permissions based on evidence.
What should your logging, monitoring and backup routine look like?
An agent makes decisions you do not watch one by one. So you need to read later what it did, when, and with which input. n8n keeps execution history in its own database, and docker compose logs shows container output. Also set retention limits and Docker log driver options per the official docs, so logs do not fill the disk.
For backups, cover three parts together:
- Data volume: The
n8n_datavolume holds the SQLite database and the encryption key. - Configuration: The Compose file, reverse proxy config and the
.envfile. Store the last one in an encrypted location. - Workflows: Exporting workflows with the n8n command line export commands helps with version control.
Do not keep backups on the same server, because if the server goes, the backup goes too. For backup frequency and retention, see our website backup strategy guide. In addition, test restores on a schedule. You cannot trust a backup you have never restored.
How do you update safely?
The n8n docs say the project releases new versions frequently and recommend the stable channel for production. That pace is good news, because security fixes ship quickly. On the other hand, updating without reading the release notes can break a working workflow.
A safe update order looks like this. First, take a backup. Then check the release notes for breaking changes. After that, change the version tag in the Compose file and run these commands:
docker compose pull
docker compose up -d
docker compose logs -f n8n
After the update, trigger your critical workflows by hand once and check the results. For example, environment variable names can change. The n8n docs note that the old variable for the webhook URL is now deprecated in favor of a new name.
Do not forget operating system updates either. Likewise, update the Ollama image and your downloaded models with the same discipline. If something breaks, roll back to the previous tag. That is also why a pinned tag is safer than the latest tag.
What should you know about GDPR and data protection?
If the agent processes personal data, self-hosting does not free you from data protection duties. In fact, as the data controller you hold full responsibility. This section offers general information only, not legal advice.
In general, you should be able to answer these questions. Which personal data does the agent read? If you use a cloud model, which provider and which country receives that data? How long does personal data stay in the execution history? If someone asks you to delete their data, can you also delete it from the agent logs?
A local model can remove the question of sending prompt text abroad. However, the server location, the backup location and the access logs still matter. You may also need to describe AI processing in your privacy notice. We cover general website compliance steps in our GDPR compliant website guide. For an agent that handles personal data, we recommend you consult a lawyer.
What are the cost items of a self-hosted AI agent?
The cost of a self-hosted AI agent goes beyond a single server invoice. We do not quote figures, because prices vary by provider and usage. Check the current rates on each provider's own pricing page. Still, we suggest you list these items separately in your budget:
- Server: VPS or GPU server rent. With a local model, this item grows noticeably.
- Model usage: Per token fees if you call a cloud API. Set a spending cap in the provider dashboard.
- Storage and backups: External storage for your backups.
- Domain and certificate: The domain renewal fee. Let's Encrypt certificates cost nothing.
- Licensing: Check the license terms for the orchestrator features you use on its official page.
- People time: Hours for setup, updates, log review and incident response. This is often the largest item.
For instance, even if the monthly server fee looks low, you underestimate the real cost when you ignore a few hours of maintenance each month.
When should you hand the setup to experts?
Not every agent deployment is a weekend project. If any of these apply to you, we suggest an experienced team or a managed service:
- You do not know what to do if a security incident hits the server.
- The agent will touch customer data, a payment system or a bulk messaging channel.
- Nobody could take over the system if the person who built it leaves.
- You have no time for weekly updates and log checks.
In these cases you have two paths. The first is the orchestrator's own cloud service, where the provider handles maintenance. The second is to hand setup and upkeep to a team. Through our AI automation services, we focus on workflow design and safe configuration, and we work with your hosting provider on the server side.
Whichever path you choose, write down what the agent can access and which actions need approval. Then the limits stay the same no matter who runs the system.
A short checklist for your first week
Once setup is complete, we recommend these checks during the first week. The list sums up the steps above:
- Confirm that only SSH, HTTP and HTTPS ports are open from outside.
- Test that nobody can reach the n8n and Ollama ports from the internet.
- Keep a safe copy of the encryption key and the
.envfile off the server. - Take your first backup and try a restore on another machine.
- Review the permissions of every integration and remove write access you do not need.
- Make sure every irreversible action has an approval step.
- Turn on a spending cap and alerts with your model provider.
In short, running an agent on your own server takes only a few commands. The real work is keeping that agent secure, backed up and limited. The official n8n Docker Compose docs, the Ollama FAQ and the Docker packet filtering and firewall docs are the core references along the way.



