{"id":22015,"date":"2026-06-15T14:02:55","date_gmt":"2026-06-15T08:32:55","guid":{"rendered":"https:\/\/www.youstable.com\/blog\/?p=22015"},"modified":"2026-06-15T14:16:57","modified_gmt":"2026-06-15T08:46:57","slug":"how-to-install-and-run-ollama-on-a-vps","status":"publish","type":"post","link":"https:\/\/www.youstable.com\/blog\/how-to-install-and-run-ollama-on-a-vps\/","title":{"rendered":"How to Install and Run Ollama on a VPS (2026 Guide)"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Yes, you can run Ollama on a VPS, and I have done it myself, multiple times, across different VPS configurations.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I have personally tested every command in this guide on a live Ubuntu 24.04 VPS, pulled models ranging from 1B to 70B and verified the API, Nginx proxy and security setup.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I have added notes at every step to make sure you do not miss a thing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You already have a VPS and you want to run AI models on it and you do not want to send your data to a third-party API.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Warning:<\/strong> Please do not think of setting up and running Ollama on your local computer. It is really a big disaster. Believe me! Your system will suffer back to back lagging and may even crash at times. That\u2019s why, no PC.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">I will tell you every step. I will tell you how to connect to your VPS, install Ollama, pull an AI model, run it, and expose it safely over the internet.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">By the time you finish this guide, your VPS will be serving a live AI model API you can call from anywhere. And in doing all that:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">NO GPU is required at all. A CPU-only VPS with 4 GB of RAM is enough to get started.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"what-is-ollama-and-why-run-ollama-on-vps\">What Is Ollama &amp; Why Run Ollama on VPS?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Ollama is an open-source runtime that lets you download and run large language models (LLMs) on your own hardware using a single command.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I have been working with VPS infrastructure for years and Ollama is the cleanest self-hosted LLM solution I\u2019ve come across. It handles model downloads, runtime management and serves an OpenAI-compatible REST API as well.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Popular models you can run with Ollama include:<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Llama 3\u00a0<\/li>\n\n\n\n<li>Mistral\u00a0<\/li>\n\n\n\n<li>Gemma<\/li>\n\n\n\n<li>DeepSeek\u00a0<\/li>\n\n\n\n<li>Qwen<\/li>\n\n\n\n<li>Phi<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">All of them run locally and your prompts and responses never leave your server.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But if you CAN run Ollama locally on your computer (PLEASE DO THAT ON YOUR RISK), then why to run Ollama on a VPS? Why is a VPS better than your local machine?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">See! The thing is that running Ollama locally is fine for quick experiments.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The moment you want to keep the model running 24\/7 or you want to access it from your phone or you want to share it with a teammate without exposing your home IP, your local machine stops making sense and a VPS becomes the only choice.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You can even integrate Ollama into a web app or automation workflow and also keep sensitive data off third-party APIs when you run Ollama on VPS.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So here is what you gain the moment you move Ollama to a VPS. Every one of these I have verified in testing environments.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Zero API costs: No per-token billing, at all<\/li>\n\n\n\n<li>Full data privacy: Prompts and responses stay on your server<\/li>\n\n\n\n<li>Always-on availability: 24\/7 uptime, no sleep mode<\/li>\n\n\n\n<li>Remote access: Call your model from any device, anywhere<\/li>\n\n\n\n<li>Scalable: upgrade RAM and CPU in just one click as soon as the need grows<\/li>\n<\/ul>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"vps-requirement-for-ollama-what-you-need-before-you-start\">VPS Requirement for Ollama: What You Need Before You Start<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Before you type a single command, make sure your VPS meets the minimum requirements. I have run Ollama on both low-end and high-end VPS instances.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Believe me! The biggest problem is RAM. The entire model must fit in memory or speed drops really very badly.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"minimum-hardware-requirements\">Minimum Hardware Requirements<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Use this table to check your VPS against the minimum and recommended specs before you begin.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td class=\"has-text-align-center\" data-align=\"center\"><strong>Spec<\/strong><\/td><td class=\"has-text-align-center\" data-align=\"center\"><strong>Minimum&nbsp;<\/strong><\/td><td class=\"has-text-align-center\" data-align=\"center\"><strong>Recommended<\/strong><\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">CPU<\/td><td class=\"has-text-align-center\" data-align=\"center\">2 vCPU<\/td><td class=\"has-text-align-center\" data-align=\"center\">4 vCPU or more<\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">RAM<\/td><td class=\"has-text-align-center\" data-align=\"center\">4 GB<\/td><td class=\"has-text-align-center\" data-align=\"center\">8 GB or more<\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">Disk<\/td><td class=\"has-text-align-center\" data-align=\"center\">20 GB free<\/td><td class=\"has-text-align-center\" data-align=\"center\">40 GB or more<\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">OS<\/td><td class=\"has-text-align-center\" data-align=\"center\">Ubuntu 22.04<\/td><td class=\"has-text-align-center\" data-align=\"center\">Ubuntu 24.04 LTS<\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">GPU<\/td><td class=\"has-text-align-center\" data-align=\"center\">Not required<\/td><td class=\"has-text-align-center\" data-align=\"center\">NVIDIA (optional, for speed)<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"how-much-ram-does-each-model-need\">How Much RAM Does Each Model Need?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The table below shows approximate RAM requirements for quantized Q4 models. Full-precision models need two to three times more RAM than these figures.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Start with a 3B or 7B model.<\/strong><\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td class=\"has-text-align-center\" data-align=\"center\"><strong>LARGE LANGUAGE MODELS<\/strong><\/td><td class=\"has-text-align-center\" data-align=\"center\"><strong>Parameters<\/strong><\/td><td class=\"has-text-align-center\" data-align=\"center\"><strong>Minimum RAM<\/strong><\/td><td class=\"has-text-align-center\" data-align=\"center\"><strong>Comfortable RAM<\/strong><\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">Llama 3.2 1B, Phi-3 mini<\/td><td class=\"has-text-align-center\" data-align=\"center\">1-3B<\/td><td class=\"has-text-align-center\" data-align=\"center\">4 GB<\/td><td class=\"has-text-align-center\" data-align=\"center\">8 GB<\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">Llama 3.1, Mistral 7B, Gemma 3 8B<\/td><td class=\"has-text-align-center\" data-align=\"center\">7-8B<\/td><td class=\"has-text-align-center\" data-align=\"center\">8 GB<\/td><td class=\"has-text-align-center\" data-align=\"center\">16 GB<\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">Llama 3.1 13B, Gemma 3 12B<\/td><td class=\"has-text-align-center\" data-align=\"center\">13B<\/td><td class=\"has-text-align-center\" data-align=\"center\">16 GB<\/td><td class=\"has-text-align-center\" data-align=\"center\">32 GB<\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">Qwen 2.5 32B, DeepSeek R1 32B<\/td><td class=\"has-text-align-center\" data-align=\"center\">32B<\/td><td class=\"has-text-align-center\" data-align=\"center\">32 GB<\/td><td class=\"has-text-align-center\" data-align=\"center\">64 GB<\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">Llama 3.3 70B, DeepSeek R1 70B<\/td><td class=\"has-text-align-center\" data-align=\"center\">70B<\/td><td class=\"has-text-align-center\" data-align=\"center\">64 GB<\/td><td class=\"has-text-align-center\" data-align=\"center\">128 GB<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">And, what about the operating system? The OS? Ubuntu 22.04 LTS or Ubuntu 24.04 LTS is the recommended OS. The Ollama install script handles everything automatically on both versions.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">All commands in this guide assume you are running Ubuntu. With your VPS ready and SSH credentials in hand, let&#8217;s go step by step.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"step-zero-get-your-vps-from-youstable\">Step Zero: Get Your VPS from YouStable<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Before you run a single command in this guide, you need a VPS. I recommend YouStable and here is exactly why it fits Ollama perfectly.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td>YouStable VPS plans are built on KVM architecture with NVMe SSD storage and guaranteed dedicated resources, which means no RAM or CPU sharing with other users. That matters a lot when you are loading a 4 GB language model into memory.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Every plan includes Ubuntu 24.04 LTS support (required for this guide), full root SSH access, advanced DDoS protection, free migration, and 99.95% uptime, exactly what you need for a 24\/7 Ollama deployment.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Now, Let\u2019s discuss how to buy a YouStable VPS: Step by Step<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Go to the official YouStable VPS page >> <a href=\"http:\/\/youstable.com\/vps-hosting\" target=\"_blank\" rel=\"noopener\"><strong>youstable.com\/vps-hosting<\/strong><\/a><\/li>\n<\/ul>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><a href=\"https:\/\/www.youstable.com\/vps-hosting\/\"><img loading=\"lazy\" decoding=\"async\" width=\"1516\" height=\"777\" src=\"https:\/\/www.youstable.com\/blog\/wp-content\/uploads\/2026\/06\/YouStable-VPS-Hosting.jpg\" alt=\"YouStable VPS Hosting\" class=\"wp-image-22029\"\/><\/a><\/figure>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Choose your plan based on the RAM table above >> For beginners, start with vStart (4 GB) or vPro (6 GB) >> Click &#8220;Buy Now&#8221; on your chosen plan.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Select your billing cycle (Choose Annual as it saves more).<\/strong><\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><a href=\"https:\/\/www.youstable.com\/vps-hosting\/\"><img loading=\"lazy\" decoding=\"async\" width=\"1546\" height=\"770\" src=\"https:\/\/www.youstable.com\/blog\/wp-content\/uploads\/2026\/06\/VPS-Hosting-Plans-Pricing.jpg\" alt=\"VPS Hosting Plans &amp; Pricing\" class=\"wp-image-22030\"\/><\/a><\/figure>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Choose your data center location >> Select USA for global access or India for Indian traffic >> Select Ubuntu 24.04 LTS as your operating system.<\/li>\n<\/ul>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><a href=\"https:\/\/www.youstable.com\/vps-hosting\/\"><img loading=\"lazy\" decoding=\"async\" width=\"1663\" height=\"2048\" src=\"https:\/\/www.youstable.com\/blog\/wp-content\/uploads\/2026\/06\/557.jpg\" alt=\"\" class=\"wp-image-22031\"><\/a><\/figure>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Complete checkout >> Your VPS login credentials and IP address will be emailed to you.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Use those credentials to SSH into your VPS and follow this guide from Step 1.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"step-1-connect-to-your-vps-via-ssh\">Step 1: Connect to Your VPS via SSH<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Everything you do with Ollama is done through the terminal. You connect to your VPS using SSH. This is how every VPS is managed, and it is the first skill you need to have before anything else.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"what-is-ssh-and-why-you-need-it\">What Is SSH and Why You Need It<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">SSH gives you a secure, encrypted connection from your local computer to your remote VPS. Once connected, you see a command prompt that controls the remote server, exactly as if you were sitting in front of it.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"how-to-connect-linux-and-macos\">How to Connect: Linux and macOS<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Open your terminal and run the following command. Replace YOUR_VPS_IP with the IP address from your VPS control panel.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>ssh root@YOUR_VPS_IP<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">If your provider gave you a non-root user, which is common on Ubuntu, use that username instead:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>ssh ubuntu@YOUR_VPS_IP<\/strong><\/code><\/pre>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"how-to-connect-windows\">How to Connect: Windows<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">On Windows 10 or later, open PowerShell or Command Prompt and run the same ssh command above. You can also use PuTTY: enter your VPS IP in the Host Name field, select port 22, and click Open.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>First Login: Accept the Host Key<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The first time you connect, you will see a message asking you to confirm the host fingerprint. Type yes and press Enter. This is normal and expected.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1055\" height=\"711\" src=\"https:\/\/www.youstable.com\/blog\/wp-content\/uploads\/2026\/06\/558.jpg\" alt=\"\" class=\"wp-image-22032\"><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The authenticity of host &#8216;1.2.3.4&#8217; can&#8217;t be established.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Are you sure you want to continue connecting (yes\/no)?<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Note<\/strong>: Your machine saves the server&#8217;s fingerprint after you type yes. On all future connections it verifies this fingerprint silently. This is a security feature, not an error<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"step-2-update-your-vps-system\">Step 2: Update Your VPS System<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Before installing anything, update the system package list. I always do this first on every fresh VPS. It ensures you are installing the latest, most secure versions of all dependencies and avoids version conflicts later.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>apt update &amp;&amp; apt upgrade -y<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">This command does two things: apt update fetches the latest list of available packages from Ubuntu&#8217;s repositories, and apt upgrade -y installs all pending updates automatically. The -y flag confirms all prompts without asking you.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1071\" height=\"715\" src=\"https:\/\/www.youstable.com\/blog\/wp-content\/uploads\/2026\/06\/559.jpg\" alt=\"\" class=\"wp-image-22033\"><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">This step typically takes one to two minutes. Wait for it to finish before proceeding.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"step-3-install-ollama-on-your-vps\">Step 3: Install Ollama on Your VPS<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Ollama provides an official installer script that handles everything.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It detects your OS, downloads the correct binary, creates an ollama system user, and sets up a systemd service so Ollama starts automatically on boot. I have run this script on dozens of VPS instances and it has never failed on a clean Ubuntu install.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Run the Install Script<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Run this single command. That is all it takes to get Ollama fully installed and running as a background service.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>curl -fsSL https:\/\/ollama.com\/install.sh | sh<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The script downloads the binary from the official Ollama release page and installs it to \/usr\/local\/bin\/ollama. It also configures Ollama to run as a background service managed by systemd.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Note<\/strong>: If your VPS has no GPU, you will see: &#8216;No NVIDIA\/AMD GPU detected. Ollama will run in CPU-only mode.&#8217; This is expected and completely fine. Your models will still run using the CPU.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1073\" height=\"690\" src=\"https:\/\/www.youstable.com\/blog\/wp-content\/uploads\/2026\/06\/561.jpg\" alt=\"\" class=\"wp-image-22034\"><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Verify the Installation<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Check that the Ollama service is running after the install script completes. This is the first thing I verify on every setup.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>systemctl status ollama<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">You should see active (running) in the output. This means Ollama is running in the background and listening on port 11434 on localhost. You can also verify the installed version:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>ollama --version<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">If both commands return output without errors, Ollama is successfully installed on your VPS.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"step-4-pull-your-first-ai-model\">Step 4: Pull Your First AI Model<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">After installation, Ollama has no models downloaded. You need to pull one before you can use it. Models are pulled from the official Ollama library at ollama.com\/library.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I recommend starting with a small model to verify everything works before pulling larger ones.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Recommended Models for Beginners on VPS<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For a VPS with 4 to 8 GB of RAM and no GPU, these are the models I personally recommend starting with. Each one has been tested and runs well on entry-level hardware.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td class=\"has-text-align-center\" data-align=\"center\"><strong>LLMs<\/strong><\/td><td class=\"has-text-align-center\" data-align=\"center\"><strong>Size<\/strong><\/td><td class=\"has-text-align-center\" data-align=\"center\"><strong>Best For<\/strong><\/td><td class=\"has-text-align-center\" data-align=\"center\"><strong>Pull Command<\/strong><\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">Llama 3.2 3B<\/td><td class=\"has-text-align-center\" data-align=\"center\">~2 GB<\/td><td class=\"has-text-align-center\" data-align=\"center\">General chat, fast responses<\/td><td class=\"has-text-align-center\" data-align=\"center\">ollama pull llama3.2:3b<\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">Qwen 2.5 3B<\/td><td class=\"has-text-align-center\" data-align=\"center\">~1.9 GB<\/td><td class=\"has-text-align-center\" data-align=\"center\">Multilingual, general use<\/td><td class=\"has-text-align-center\" data-align=\"center\">ollama pull qwen2.5:3b<\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">Phi-3 Mini<\/td><td class=\"has-text-align-center\" data-align=\"center\">~2.3 GB<\/td><td class=\"has-text-align-center\" data-align=\"center\">Reasoning, low RAM usage<\/td><td class=\"has-text-align-center\" data-align=\"center\">ollama pull phi3:mini<\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">Mistral 7B<\/td><td class=\"has-text-align-center\" data-align=\"center\">~4.1 GB<\/td><td class=\"has-text-align-center\" data-align=\"center\">General purpose (needs 8 GB RAM)<\/td><td class=\"has-text-align-center\" data-align=\"center\">ollama pull mistral:7b<\/td><\/tr><tr><td class=\"has-text-align-center\" data-align=\"center\">DeepSeek R1 8B<\/td><td class=\"has-text-align-center\" data-align=\"center\">~4.9 GB<\/td><td class=\"has-text-align-center\" data-align=\"center\">Structured reasoning<\/td><td class=\"has-text-align-center\" data-align=\"center\">ollama pull deepseek-r1:8b<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>How to Pull a Model<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Run the pull command for your chosen model. Ollama downloads the model file directly to your VPS. The download size ranges from 1.5 GB to 5 GB depending on the model.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>ollama pull llama3.2:3b<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Wait for the progress bar to complete before running the model. Do not interrupt the download.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1059\" height=\"679\" src=\"https:\/\/www.youstable.com\/blog\/wp-content\/uploads\/2026\/06\/562.jpg\" alt=\"\" class=\"wp-image-22035\"><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Manage Your Downloaded Models<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">These two commands let you see what is downloaded and remove models to free disk space when needed.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>ollama list\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 (# list all downloaded models)ollama rm llama3.2:3b \u00a0 (# remove a specific model)<\/strong><\/code><\/pre>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"step-5-run-the-model-and-test-it\">Step 5: Run the Model and Test It<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Once the model is downloaded, you can start using it immediately.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Ollama gives you two ways to interact with it: a terminal chat session and a REST API. I test both every time I set up a new instance to confirm everything is wired up correctly.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Terminal Chat Mode<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">To open an interactive chat session in your terminal, run the command below. You will see a prompt where you can type and receive responses directly in the terminal.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>ollama run llama3.2:3b<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">You will see a prompt like &gt;&gt;&gt; . Type your message and press Enter. To exit the chat, type \/bye or press Ctrl+D.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td>Note: On a CPU-only VPS with a 3B model, expect around 5 to 10 tokens per second. This is perfectly usable for most tasks. Larger models or a GPU will increase this speed significantly.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1068\" height=\"713\" src=\"https:\/\/www.youstable.com\/blog\/wp-content\/uploads\/2026\/06\/563.jpg\" alt=\"\" class=\"wp-image-22036\"><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Test the REST API<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Ollama automatically starts an API server on http:\/\/localhost:11434. Test it without leaving the terminal using curl. This is the command I use every time to confirm the API is working.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>curl http:\/\/localhost:11434\/api\/generate -d '{\"model\": \"llama3.2:3b\", \"prompt\": \"What is a VPS?\", \"stream\": false}'<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">You will get a JSON response with the model&#8217;s reply. If you see a JSON object with a response field, everything is working correctly.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>curl http:\/\/localhost:11434<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">This should return Ollama is running. That confirms the API endpoint is active and reachable from within the VPS.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"step-6-keep-ollama-running-as-a-background-service\">Step 6: Keep Ollama Running as a Background Service<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Ollama is already set up as a systemd service during installation. This means it starts automatically when your VPS reboots. Here are the key commands to manage it. I keep these in my notes for every deployment.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Service Management Commands<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Use these commands to start, stop, restart, and monitor the Ollama service on your VPS at any time.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>sudo systemctl status ollama<\/strong> # check if running<strong>sudo systemctl start ollama<\/strong> # start the service<strong>sudo systemctl stop ollama<\/strong>\u00a0 # stop the service<strong>sudo systemctl restart ollama <\/strong>\u00a0\u00a0# restart after config changes<strong>sudo systemctl enable ollama<\/strong> # start automatically on boot<\/code><\/pre>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1067\" height=\"723\" src=\"https:\/\/www.youstable.com\/blog\/wp-content\/uploads\/2026\/06\/564.jpg\" alt=\"\" class=\"wp-image-22037\"><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Check Logs<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If Ollama is not behaving as expected, check the service logs. The -f flag streams new log entries in real time so you can watch what is happening live.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>journalctl -u ollama -f<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Press Ctrl+C to stop watching the log stream.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"step-7-enable-remote-access-to-the-ollama-api\">Step 7: Enable Remote Access to the Ollama API<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">By default, Ollama only accepts connections from localhost. This means you can only call the API from within the VPS itself. To access it from your laptop, browser, or another server, you need to expose the API safely.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I always use the Nginx reverse proxy approach in production.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"option-a-bind-ollama-to-all-interfaces-simple-but-less-secure\">Option A: Bind Ollama to All Interfaces (Simple but Less Secure)<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">This makes Ollama directly accessible on port 11434 from any IP. Only use this on a private network or for testing purposes when you fully understand the risks.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>sudo systemctl edit ollama<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Add these lines in the editor:<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">[Service]\n\n\n\n<p class=\"wp-block-paragraph\">Environment=&#8221;OLLAMA_HOST=0.0.0.0:11434&#8243;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Save the file, then reload and restart:<\/strong><\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>sudo systemctl daemon-reloadsudo systemctl restart ollama<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Now open port 11434 in your VPS firewall:<\/strong><\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>sudo ufw allow 11434\/tcp<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">You can now call the API from outside at http:\/\/YOUR_VPS_IP:11434<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td>Disclaimer: Never leave port 11434 open to the public internet without authentication. This option is for testing only. Use Option B for any production or semi-public deployment.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"option-b-nginx-reverse-proxy-with-https-recommended\">Option B: Nginx Reverse Proxy with HTTPS (Recommended)<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">This is the proper production approach. Keep Ollama on localhost and put Nginx in front of it to handle HTTPS, authentication, and rate limiting. This is what I use on every production deployment.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Install Nginx and Certbot<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Run this command to install Nginx and the Certbot plugin for automatic SSL certificate management.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>sudo apt install nginx certbot python3-certbot-nginx -y<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Create the Nginx Site Configuration. Create a new config file. Replace api.yourdomain.com with your actual domain name throughout this configuration.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>sudo nano \/etc\/nginx\/sites-available\/ollama<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Paste the following configuration into the file:<\/strong><\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>server {\u00a0\u00a0\u00a0\u00a0listen 80;\u00a0\u00a0\u00a0\u00a0server_name api.yourdomain.com;\u00a0\u00a0\u00a0\u00a0\u00a0location \/ {\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0proxy_pass http:\/\/127.0.0.1:11434;\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0proxy_set_header Host $host;\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0proxy_set_header X-Real-IP $remote_addr;\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0proxy_set_header X-Forwarded-Proto $scheme;\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0proxy_http_version 1.1;\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0proxy_set_header Upgrade $http_upgrade;\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0proxy_set_header Connection \"upgrade\";\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0proxy_buffering off;\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0proxy_cache off;\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0proxy_read_timeout 600s;\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0proxy_send_timeout 600s;\u00a0\u00a0\u00a0\u00a0}}<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Enable the site and test the configuration:<\/strong><\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>sudo ln -s \/etc\/nginx\/sites-available\/ollama \/etc\/nginx\/sites-enabled\/sudo nginx -tsudo systemctl reload nginx<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\u00a0Add SSL with Certbot<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Get a free SSL certificate from Let&#8217;s Encrypt. Make sure your domain&#8217;s DNS A record points to your VPS IP address before running this command.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>sudo certbot --nginx -d api.yourdomain.com<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Follow the prompts. Certbot will automatically update your Nginx config to redirect HTTP to HTTPS and install the certificate. Your Ollama API will now be available at https:\/\/api.yourdomain.com.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"step-8-secure-your-ollama-deployment\">Step 8: Secure Your Ollama Deployment<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A default Ollama installation exposed to the public internet is a serious security risk. I have seen this firsthand: according to Shodan, thousands of Ollama servers worldwide have port 11434 open with no authentication.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Anyone who finds your endpoint can use your server resources, pull models, or flood the API. Securing your setup takes five minutes.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Configure UFW Firewall. Set up UFW to block port 11434 from the outside while allowing SSH, HTTP, and HTTPS traffic through.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Always run these commands in this exact order.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>sudo ufw allow 22\/tcpsudo ufw allow 80\/tcpsudo ufw allow 443\/tcpsudo ufw deny 11434sudo ufw enablesudo ufw status<\/strong><\/code><\/pre>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Disclaimer<\/strong>: Always allow port 22 (SSH) BEFORE enabling UFW. If you enable UFW without allowing SSH first, you will lock yourself out of the server completely.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Add Basic Authentication to Nginx<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Install htpasswd and create a password file. Then add authentication to your Nginx config so every request to your Ollama API requires a username and password.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>sudo apt install apache2-utils -ysudo htpasswd -c \/etc\/nginx\/.ollama_passwd your_username<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Add these two lines inside the location block in your Nginx config:<\/strong><\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>auth_basic \"Restricted\";auth_basic_user_file \/etc\/nginx\/.ollama_passwd;<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Reload Nginx to apply the changes:<\/strong><\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>sudo systemctl reload nginx<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Security Checklist<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Go through this checklist before you consider your Ollama deployment production-ready. I verify every item on this list before handing off any deployment.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>UFW enabled with port 11434 blocked from outside<\/li>\n\n\n\n<li>SSH key authentication enabled (disable password login)<\/li>\n\n\n\n<li>HTTPS via Let&#8217;s Encrypt certificate in place<\/li>\n\n\n\n<li>Basic auth or API key added to Nginx<\/li>\n\n\n\n<li>Ollama bound to localhost only (OLLAMA_HOST not set to 0.0.0.0)<\/li>\n\n\n\n<li>Regular system updates applied (apt update &amp;&amp; apt upgrade)<\/li>\n<\/ul>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"install-open-webui-a-browser-interface-for-your-models\">Install Open WebUI: A Browser Interface for Your Models<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">If you prefer a visual chat interface instead of the terminal or raw API calls, Open WebUI gives you a ChatGPT-style browser interface that connects directly to your Ollama server.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It runs in Docker and takes about five minutes to set up. I use it on every internal deployment.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Install Docker &gt;&gt; Run these three commands to install Docker and configure it to start automatically with the system.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>curl -fsSL https:\/\/get.docker.com | shsudo systemctl start dockersudo systemctl enable docker<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Run Open WebUI<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Run the following Docker command to pull and start Open WebUI. The &#8211;restart always flag ensures it comes back up automatically after a reboot.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>docker run -d -p 3000:8080 \\\u00a0\u00a0--add-host=host.docker.internal:host-gateway \\\u00a0\u00a0-v open-webui:\/app\/backend\/data \\\u00a0\u00a0--name open-webui \\\u00a0\u00a0--restart always \\\u00a0\u00a0ghcr.io\/open-webui\/open-webui:main<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Open WebUI will be accessible at http:\/\/YOUR_VPS_IP:3000.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On your first visit, create an admin account. It will automatically connect to your Ollama instance and show all your downloaded models.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Note<\/strong>: You can also put Open WebUI behind Nginx with HTTPS following the same reverse proxy steps from Section 9. Just create a second server block pointing to port 3000.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"troubleshooting-common-issues\">Troubleshooting Common Issues<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Here are the most common problems people run into when setting up Ollama on a VPS, and exactly how to fix each one. I have encountered and resolved every issue listed here during my own deployments.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">&#8220;Connection refused&#8221; When Calling the API<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This usually means Ollama is not running. Check the service status and start it if it shows inactive or failed.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>systemctl status ollamasudo systemctl start ollama<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>If it keeps failing, check the last 50 lines of logs for the actual error:<\/strong><\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>journalctl -u ollama --no-pager -n 50<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Model Is Very Slow<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On a CPU-only VPS, LLM generation is slower than GPU-accelerated setups. Typical speeds are 3 to 10 tokens per second for a 3B to 7B model. These steps help maximize your available performance.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Use a smaller model: try llama3.2:1b or qwen2.5:3b instead of 7B models<\/li>\n\n\n\n<li>Check if the model is fully in RAM: run &#8220;free -h&#8221; and compare available RAM to model size<\/li>\n\n\n\n<li>Close other memory-heavy processes to free up RAM<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">&#8220;Error: model not found&#8221;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You need to pull the model before you can run it. Use ollama list to see all models that are currently downloaded.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>ollama pull MODEL_NAMEollama list<\/strong><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Nginx Returns 502 Bad Gateway<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This means Nginx is running but cannot reach Ollama. Work through this checklist in order until the issue is resolved.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Ollama is running: systemctl status ollama<\/li>\n\n\n\n<li>OLLAMA_HOST is not set to 0.0.0.0 when using Nginx proxy (keep it on localhost)<\/li>\n\n\n\n<li>proxy_pass in Nginx points to http:\/\/127.0.0.1:11434<\/li>\n\n\n\n<li>Nginx config is valid: sudo nginx -t<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">VPS Runs Out of RAM While Loading Model<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The model size exceeds your available RAM. Switch to a smaller model, upgrade your VPS RAM tier, or add swap space as a temporary buffer.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>sudo fallocate -l 4G \/swapfilesudo chmod 600 \/swapfilesudo mkswap \/swapfilesudo swapon \/swapfile<\/strong><\/code><\/pre>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td>Note: Swap space is a short-term workaround only. It allows your VPS to use disk as overflow RAM, but disk is much slower than actual RAM. For sustained performance, upgrade to a VPS with more RAM.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"use-cases-what-you-can-build-with-ollama-on-vps\">Use Cases: What You Can Build With Ollama on VPS<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Once your Ollama instance is running on a VPS with a stable API endpoint, you can integrate it into a wide range of real-world applications. These are the use cases I see most often in production environments.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"private-ai-chatbot-for-your-team\">Private AI Chatbot for Your Team<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Use Open WebUI on top of Ollama to give your team a private, self-hosted chat interface. No data ever leaves your server, which makes it suitable for legal, medical, financial, or confidential internal content.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"automated-workflows-with-n8n\">Automated Workflows with n8n<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">n8n is a popular open-source automation tool. Connect it to your Ollama API to build workflows that summarize emails, extract data from documents, generate content, or reply to support tickets, all without paying per API call.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"rag-pipeline-over-internal-documents\">RAG Pipeline Over Internal Documents<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Combine Ollama with a vector database like Chroma or Qdrant to build a retrieval-augmented generation system. Your Ollama model can answer questions based on your internal documents, knowledgebase, or codebase.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"api-backend-for-your-web-app\">API Backend for Your Web App<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Because Ollama exposes an OpenAI-compatible REST API, any application that already integrates with OpenAI can be pointed to your Ollama endpoint instead with no code changes. Just update the base URL and eliminate your OpenAI bill.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"ai-powered-dev-tools\">AI-Powered Dev Tools<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Use Ollama with coding-focused models like qwen2.5-coder or deepseek-coder to power local code completion, review automation, or documentation generation inside your CI\/CD pipeline.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"quick-reference-all-commands-in-one-place\">Quick Reference: All Commands in One Place<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Here is every command from this guide grouped by task for easy reference. Bookmark this section so you do not have to scroll through the full guide every time you need a command.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"installation\">Installation<\/h3>\n\n\n\n<pre class=\"wp-block-code\"><code>ssh root@YOUR_VPS_IPapt update &amp;&amp; apt upgrade -ycurl -fsSL https:\/\/ollama.com\/install.sh | shsystemctl status ollama<\/code><\/pre>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"model-management\">Model Management<\/h3>\n\n\n\n<pre class=\"wp-block-code\"><code>ollama pull llama3.2:3bollama run llama3.2:3bollama listollama rm MODEL_NAME<\/code><\/pre>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"api-test\">API Test<\/h3>\n\n\n\n<pre class=\"wp-block-code\"><code>curl http:\/\/localhost:11434curl http:\/\/localhost:11434\/api\/generate -d '{\"model\":\"llama3.2:3b\",\"prompt\":\"Hello\",\"stream\":false}'<\/code><\/pre>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"service-management\">Service Management<\/h3>\n\n\n\n<pre class=\"wp-block-code\"><code>sudo systemctl start ollamasudo systemctl stop ollamasudo systemctl restart ollamasudo systemctl enable ollamajournalctl -u ollama -f<\/code><\/pre>\n\n\n\n<h3 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"firewall\">Firewall<\/h3>\n\n\n\n<pre class=\"wp-block-code\"><code>sudo ufw allow 22\/tcpsudo ufw allow 80\/tcpsudo ufw allow 443\/tcpsudo ufw deny 11434sudo ufw enable<\/code><\/pre>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"faqs\">FAQs<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">These are the questions I get asked most often by people setting up Ollama on a VPS for the first time.<\/p>\n\n\n<div id=\"rank-math-faq\" class=\"rank-math-block\">\n<div class=\"rank-math-list \">\n<div id=\"faq-question-1781505865970\" class=\"rank-math-list-item\">\n<h3 class=\"rank-math-question \" class=\"rank-math-question \" id=\"do-i-need-a-gpu-to-run-ollama-on-a-vps\">Do I need a GPU to run Ollama on a VPS?<\/h3>\n<div class=\"rank-math-answer \">\n\n<p>No. Ollama runs on CPU-only VPS instances without any issue. For 1B to 7B models, a CPU-only setup with 4 to 8 GB of RAM gives acceptable performance. GPU acceleration is optional but makes generation 5 to 10 times faster.<\/p>\n\n<\/div>\n<\/div>\n<div id=\"faq-question-1781505884756\" class=\"rank-math-list-item\">\n<h3 class=\"rank-math-question \" class=\"rank-math-question \" id=\"what-is-the-minimum-vps-size-i-need\">What is the minimum VPS size I need?<\/h3>\n<div class=\"rank-math-answer \">\n\n<p>A VPS with 2 vCPU and 4 GB of RAM is the absolute minimum. It will run 1B to 3B models comfortably. For 7B models, use at least 8 GB of RAM.<\/p>\n\n<\/div>\n<\/div>\n<div id=\"faq-question-1781505894632\" class=\"rank-math-list-item\">\n<h3 class=\"rank-math-question \" class=\"rank-math-question \" id=\"is-ollama-free-to-use-on-a-vps\">Is Ollama free to use on a VPS?<\/h3>\n<div class=\"rank-math-answer \">\n\n<p>Yes. Ollama is free and open-source. The only cost is your <strong><a href=\"https:\/\/www.youstable.com\/vps-hosting\/\">VPS hosting<\/a><\/strong> fee. You pay nothing per model download or per API call.<\/p>\n\n<\/div>\n<\/div>\n<div id=\"faq-question-1781505902257\" class=\"rank-math-list-item\">\n<h3 class=\"rank-math-question \" class=\"rank-math-question \" id=\"can-i-run-multiple-models-on-the-same-vps\">Can I run multiple models on the same VPS?<\/h3>\n<div class=\"rank-math-answer \">\n\n<p>Yes. You can have multiple models downloaded at once. Ollama only loads a model into memory when you call it. Switch between models simply by changing the model name in your API request.<\/p>\n\n<\/div>\n<\/div>\n<div id=\"faq-question-1781505909131\" class=\"rank-math-list-item\">\n<h3 class=\"rank-math-question \" class=\"rank-math-question \" id=\"is-port-11434-safe-to-leave-open\">Is port 11434 safe to leave open?<\/h3>\n<div class=\"rank-math-answer \">\n\n<p>No. Never expose port 11434 directly to the public internet without authentication. Use an Nginx reverse proxy with HTTPS and basic auth, and block port 11434 with UFW as described in this guide.<\/p>\n\n<\/div>\n<\/div>\n<div id=\"faq-question-1781505917888\" class=\"rank-math-list-item\">\n<h3 class=\"rank-math-question \" class=\"rank-math-question \" id=\"can-i-use-ollama-with-openai-compatible-libraries\">Can I use Ollama with OpenAI-compatible libraries?<\/h3>\n<div class=\"rank-math-answer \">\n\n<p>Yes. Ollama exposes an OpenAI-compatible REST API at \/v1\/chat\/completions. Any SDK or library that supports a custom base URL, including the official OpenAI Python SDK, works with Ollama without code changes.<\/p>\n\n<\/div>\n<\/div>\n<div id=\"faq-question-1781505924530\" class=\"rank-math-list-item\">\n<h3 class=\"rank-math-question \" class=\"rank-math-question \" id=\"how-do-i-update-ollama\">How do I update Ollama?<\/h3>\n<div class=\"rank-math-answer \">\n\n<p>Run the official install script again. It will update Ollama to the latest version without affecting your downloaded models.<br \/>curl -fsSL https:\/\/ollama.com\/install.sh | sh<\/p>\n\n<\/div>\n<\/div>\n<div id=\"faq-question-1781505939255\" class=\"rank-math-list-item\">\n<h3 class=\"rank-math-question \" class=\"rank-math-question \" id=\"what-happens-if-my-vps-reboots\">What happens if my VPS reboots?<\/h3>\n<div class=\"rank-math-answer \">\n\n<p>Ollama is configured as a systemd service during installation. It will start automatically when the VPS boots. No manual intervention is needed after a reboot.<\/p>\n\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\" class=\"wp-block-heading\" id=\"a-final-note-to-the-reader\">A Final Note To The Reader<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">You have now set up a fully functional, private AI API on your own VPS.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I have walked you through every step personally, from SSH access to Nginx configuration and security hardening. Every command in this guide has been tested on live Ubuntu VPS instances.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Your VPS is now running Ollama as a persistent background service with HTTPS access, firewall rules in place, and authentication protecting the endpoint.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You own the infrastructure, the models, and the data.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Yes, you can run Ollama on a VPS, and I have done it myself, multiple times, across different VPS configurations.&nbsp; [&hellip;]<\/p>\n","protected":false},"author":21,"featured_media":22043,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"inline_featured_image":false,"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"iawp_total_views":169,"footnotes":""},"categories":[1136,2268,350],"tags":[],"class_list":["post-22015","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-tutorials","category-kb-ai-automation","category-knowledgebase"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.youstable.com\/blog\/wp-json\/wp\/v2\/posts\/22015","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.youstable.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.youstable.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.youstable.com\/blog\/wp-json\/wp\/v2\/users\/21"}],"replies":[{"embeddable":true,"href":"https:\/\/www.youstable.com\/blog\/wp-json\/wp\/v2\/comments?post=22015"}],"version-history":[{"count":0,"href":"https:\/\/www.youstable.com\/blog\/wp-json\/wp\/v2\/posts\/22015\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.youstable.com\/blog\/wp-json\/wp\/v2\/media\/22043"}],"wp:attachment":[{"href":"https:\/\/www.youstable.com\/blog\/wp-json\/wp\/v2\/media?parent=22015"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.youstable.com\/blog\/wp-json\/wp\/v2\/categories?post=22015"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.youstable.com\/blog\/wp-json\/wp\/v2\/tags?post=22015"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}