Windows AI Server Architecture: Running Unsloth Studio on the Host, Accessing It From WSL

β€œWe spent weeks trying to cram a GPU-accelerated model server into WSL. Then we realized Windows already had the GPU. We just needed a tunnel.”


The Problem We All Hit

If you’ve tried to run local AI inference from inside WSL2, you know the pain.

You’ve got this beautiful RTX 4070 Ti sitting in your Windows host. CUDA drivers installed, PyTorch recognizes it, life is good on the Windows side. Then you drop into WSL β€” where Hermes lives, where your agent workflows run, where all the actual work happens β€” and suddenly that GPU might as well be on Mars.

WSL2 can access the GPU via CUDA passthrough… sometimes. But when you’re running Unsloth Studio β€” which needs specific CUDA versions, Windows-native GPU binaries, and a whole stack of dependencies that fight with Linux package managers β€” β€œsometimes” isn’t good enough. Especially when you’re trying to serve models to a multi-agent orchestration system that doesn’t have patience for β€œlet me reinstall CUDA again.”

We tried:

  • Running Ollama inside WSL (works, but no Unsloth optimizations)
  • Docker-in-WSL with NVIDIA Container Toolkit (works, but adds a whole orchestration layer)
  • Running Unsloth directly in WSL (dependency hell, version conflicts, crying)
  • The ZimaBoard (remember that post? yeah, we remember too)

What we didn’t try β€” until we got desperate β€” was the obvious thing: put the model server on the Windows host, where the GPU is happy, and just talk to it from WSL.


The Revelation: Windows Host as AI Server

Here’s the architecture that actually works:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    Windows Host                             β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚  Unsloth Studio │◄──────►│  netsh portproxy rule    β”‚   β”‚
β”‚  β”‚  (GPU accelerated)       β”‚  0.0.0.0:11434 β†’         β”‚   β”‚
β”‚  β”‚  localhost:11434         β”‚  localhost:11434 (relay) β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                                         β”‚                   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                          β”‚ WSL2 virtual switch
                                          β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    WSL 2 (Hermes)                           β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚  Hermes Agent Framework                             β”‚   β”‚
β”‚  β”‚  β”œβ”€β”€ Seek profile  ──► HTTP request to            β”‚   β”‚
β”‚  β”‚  β”œβ”€β”€ Code profile  ──► host.internal:11434        β”‚   β”‚
β”‚  β”‚  └── Scribe profile ──► (Unsloth Studio API)      β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

The trick isn’t Docker. It isn’t a VPN. It isn’t some obscure WSL configuration. It’s netsh interface portproxy β€” a built-in Windows utility that’s been sitting there since Windows XP, quietly waiting for someone to need it.


The One Command That Makes It Work

On your Windows host (PowerShell as Administrator):

# Expose localhost:11434 (Unsloth Studio) to all interfaces so WSL can reach it
netsh interface portproxy add v4tov4 listenport=11434 listenaddress=0.0.0.0 connectport=11434 connectaddress=127.0.0.1

# Allow it through the Windows Firewall
netsh advfirewall firewall add rule name="Unsloth Studio WSL" dir=in action=allow protocol=tcp localport=11434

Now from inside WSL:

curl http://$(cat /etc/resolv.conf | grep nameserver | awk '{print $2}'):11434/api/tags

Boom. Your WSL instance is talking to Unsloth Studio on the Windows host. No Docker. No ssh tunnels. No VPN. Just two commands and a fundamental understanding that localhost in Windows and localhost in WSL are not the same place.

The IP you need is whatever’s in /etc/resolv.conf inside WSL β€” that’s the WSL virtual switch’s gateway IP pointing back at the Windows host. Usually 172.x.x.x. You can set an alias in your shell profile so you never have to think about it again:

# ~/.bashrc or ~/.zshrc
export WINDOWS_HOST=$(cat /etc/resolv.conf | grep nameserver | awk '{print $2}')

What This Unlocks: Unsloth Studio on Windows

Why Unsloth Studio specifically? Because if you’re doing local inference on consumer hardware, Unsloth is the difference between β€œit works” and β€œit works fast.”

Unsloth implements hand-optimized kernels for model inference that squeeze 2-5x speedups out of standard Transformers. It’s not a model β€” it’s an acceleration layer. When you run Qwen3 through Unsloth vs. stock Transformers, the difference is measurable in seconds per token, not milliseconds.

And here’s the thing: Unsloth Studio runs flawlessly on Windows. It’s a native Windows application with built-in model management, a chat interface, and an OpenAI-compatible local API. You download it, install it, point it at your models directory, and it just… works. Your GPU is fully utilized. Your VRAM is managed intelligently. You don’t fight with libcuda.so or LD_LIBRARY_PATH or any of the Linux GPU nonsense.

The API it exposes on :11434 is compatible with Ollama’s format. That means anything that can talk to Ollama can talk to Unsloth Studio. Hermes. OpenWebUI. Continue.dev. Anything.


A pane of glass like a window frame with a glowing chip behind it

Model Comparison: What Actually Runs Here

When we got this working, the first thing we did was load up the models we’d been failing to run elsewhere. Here’s the field report:

Model Size VRAM Speed (tok/s) Notes
Gemma 4 9B (E2B) 9B ~6.5GB 42 The star. Handles Unsloth’s E2B (extended context) variant with 256K context window. This is the model that froze our ZimaBoard.
Qwen3.5 0.6B–32B 2GB–20GB 28–55 Multiple variants. The 4B is the sweet spot for agent loops.
Llama 3.3 8B, 70B 6GB, 40GB 35, 8 8B runs warm. 70B needs a bigger card, but Unsloth’s quantization makes it possible on 24GB.
DeepSeek R1 (distill) 14B, 32B 10GB, 20GB 22, 12 Reasoning models. Slower but smarter for complex tasks.
Phi-4 14B 8GB 30 Microsoft’s latest. Surprisingly capable for its size.

The Gemma4 E2B comparison that matters: On our previous attempts (Ollama-in-WSL, Ollama-on-ZimaBoard), Gemma4 either failed to load, hallucinated tool calls, or froze at 0% context. Through Unsloth Studio on Windows with direct GPU acceleration? 42 tokens per second, 256K context, zero tool hallucinations. The difference isn’t the model. It’s the runtime.


The Architecture Note

Let’s talk about why this is the right architecture for a Windows+WSL dev setup, not just the convenient one.

Windows owns the hardware. Your GPU drivers, CUDA toolkit, and DirectML stack are designed for Windows first and Linux second. Every abstraction layer you add (WSL passthrough, Docker containers, remote boards) is a potential point of failure.

WSL owns the workflow. Your dev tools, your Hermes agents, your git repos, your Python environments β€” they’re all happier in Linux. The filesystem is faster. The package manager works. The terminal doesn’t lie to you.

Portproxy bridges the gap without bridging the OS. You’re not sharing CUDA between OSes. You’re not fighting with WSLg or X11 forwarding. You’re using the oldest networking primitive in the book: a TCP socket. Unsloth Studio speaks HTTP. Hermes speaks HTTP. The operating system they live on is irrelevant as long as packets can flow.

When we moved from β€œWSL runs everything” to β€œWindows runs GPU workloads, WSL runs agent logic,” our reliability went from ~60% (random CUDA errors, model load failures) to ~99%. The one percent is when Windows Update decides to reboot while you’re serving a model. We’re working on that.


Practical Setup: From Zero to Agent Loop

If you want to replicate this exact setup:

1. Install Unsloth Studio on Windows

  • Download from unsloth.ai
  • Install, launch, download a model (Gemma 4 9B E2B is a great start)
  • Verify the API is alive: curl http://localhost:11434/api/tags from PowerShell

2. Create the portproxy rule

# Run as Administrator
netsh interface portproxy add v4tov4 listenport=11434 listenaddress=0.0.0.0 connectport=11434 connectaddress=127.0.0.1
netsh advfirewall firewall add rule name="Unsloth Studio WSL" dir=in action=allow protocol=tcp localport=11434

3. Configure Hermes to point at it In your Hermes provider config (or .env):

OLLAMA_BASE_URL=http://172.x.x.x:11434  # your WSL gateway IP
# Or dynamically:
OLLAMA_BASE_URL=http://$(cat /etc/resolv.conf | grep nameserver | awk '{print $2}'):11434

4. Test from Hermes

hermes chat --model gemma4:latest --message "Say hello from the Windows side"

5. Run a kanban task with a local model

hermes kanban work t_xxxx --model gemma4:latest

A small robot pushing an enormous glowing plug into a socket of light

Benefits You Actually Feel

No more API bills. Once you own the inference, every agent run is free. We ran a full seek β†’ code β†’ scribe pipeline last night on Gemma4 and paid exactly $0.00.

No more context anxiety. 256K context windows mean your agent can read an entire codebase before answering. No more β€œplease summarize this in 4K tokens.”

No more hardware FOMO. Your existing GPU is probably enough. Unsloth’s optimizations mean a 12GB card can comfortably run 32B models. A 24GB card can run 70B with quantization.

No more β€œwill it work today?” Windows drivers are stable. Unsloth Studio is a native app. The portproxy rules persist across reboots. This setup has the reliability of a kitchen appliance, not a research project.

Agent loops that actually loop. Remember when our ZimaBoard froze at 0% context for two minutes? That doesn’t happen here. Gemma4 under Unsloth handles the full Hermes XML tool schema without breaking a sweat.


What About the ZimaBoard?

It’s still sitting there, warm and ready, running Ollama for casual chats. It’s not dead β€” it’s just scoped. Single-shot inference, chat interfaces, coding copilot. Perfectly respectable work for a $300 box.

But for the heavy lifting β€” the multi-agent orchestration, the 256K context research tasks, the iterative coding workflows where an agent writes files and another agent reviews them β€” the Windows host + Unsloth Studio setup is the production stack.

The ZimaBoard is our development league pitcher. Unsloth Studio on Windows is our closer.


The Honest Verdict

Approach GPU Access Context Speed Agent Loops Cost
ZimaBoard + Ollama 8GB RTX 3060 44K max 4s/tok ❌ Freezes $300 upfront
WSL + Ollama CUDA passthrough 128K 12 tok/s ⚠️ Unreliable $0
Windows Host + Unsloth + portproxy Full native 256K 42 tok/s βœ… Stable $0 ongoing
Kimi K2.6 API Their A100s 128K Instant βœ… Perfect ~$0.001/tok

For local agent work, the middle row wins. It gives you 90% of the cloud API experience at 0% of the per-token cost. The tradeoff is upfront hardware (which you probably already own) and a one-time setup (which you just read).


tl;dr for the Speedrunners

  • Run Unsloth Studio on Windows. It owns the GPU and it’s happy there.
  • Use netsh portproxy to forward :11434 to all interfaces.
  • Hit that port from WSL using the gateway IP in /etc/resolv.conf.
  • Configure Hermes’ OLLAMA_BASE_URL to point at it.
  • Run agent loops on Gemma4 E2B with 256K context for $0.
  • Thank me later.

β€œThey said local multi-agent AI was impossible on consumer hardware. They just hadn’t tried letting Windows be Windows and Linux be Linux.”

β€” Cleetus 🀑

P.S. β€” If your Windows Firewall complains about allowing port 11434, tell it Cleetus said it’s fine. That usually works. If it doesn’t, add the rule properly. Firewalls are no joke.

P.P.S. β€” 42 tokens per second on a 4070 Ti. That’s not a benchmark number β€” that’s what we measured while writing this post. The model was running, I was typing, and neither of us dropped a frame. That’s the Unsloth difference.