In this video, I’ll show you how to get access to a free cloud server with 32GB VRAM and use it to run AI models locally with Ollama.
If your own PC doesn’t have a powerful GPU, this method can be useful for experimenting with larger local AI models without buying expensive hardware. We’ll set everything up step by step, install Ollama, download and run a local AI model, and then create a Cloudflare Tunnel so you can access the Ollama server remotely.
We’ll also use Hugging Face to get the AI model and Kaggle as the cloud environment.
🔥 What You’ll Learn
How to get a free cloud environment for AI
How to use a server with 32GB VRAM
How to install Ollama
How to run AI models with Ollama
How to use Hugging Face models with Ollama
How to install Cloudflared
How to create a Cloudflare Tunnel
How to access your Ollama server remotely
How to run larger AI models without a powerful local GPU
🔗 Useful Links:
To Download Hermes Agent - https://studio.youtube.com/video/f8D9...
Hugging Face:
https://huggingface.co/
Kaggle:
https://www.kaggle.com/
Install Zstandard-
!sudo apt-get install zstd
Install Ollama -
!curl -fsSL https://ollama.com/install.sh | sh
Start Ollama - import subprocess
import time
subprocess.Popen(
["ollama", "serve"],
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL
)
time.sleep(5)
print("Ollama started")
Run the AI Model -
!ollama run hf.co/JonathanColetti/Qwen3.8-27B-Uncensored-GGUF:Q4_K_M
Install Cloudflared -
!wget -q https://github.com/cloudflare/cloudfl...
!dpkg -i cloudflared-linux-amd64.deb
Create the Cloudflare Tunnel -
import subprocess
import time
cloudflared = subprocess.Popen(
[
"cloudflared",
"tunnel",
"--url", "http://127.0.0.1:11434",
"--http-host-header", "localhost:11434"
],
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
text=True
)
time.sleep(8)
for _ in range(30):
line = cloudflared.stdout.readline()
if line:
print(line, end="")
#LocalAI #Ollama #AI #CloudGPU #HuggingFace #Kaggle #Cloudflare #OpenSourceAI #llm