Windows AI Server Architecture: Running Unsloth Studio on the Host, Accessing It From WSL
βWe spent weeks trying to cram a GPU-accelerated model server into WSL. Then we realized Windows already had the GPU. We just needed a tunnel.β
The Problem We All Hit
If youβve tried to run local AI inference from inside WSL2, you know the pain.
Youβve got this beautiful RTX 4070 Ti sitting in your Windows host. CUDA drivers installed, PyTorch recognizes it, life is good on the Windows side. Then you drop into WSL β where Hermes lives, where your agent workflows run, where all the actual work happens β and suddenly that GPU might as well be on Mars.
WSL2 can access the GPU via CUDA passthroughβ¦ sometimes. But when youβre running Unsloth Studio β which needs specific CUDA versions, Windows-native GPU binaries, and a whole stack of dependencies that fight with Linux package managers β βsometimesβ isnβt good enough. Especially when youβre trying to serve models to a multi-agent orchestration system that doesnβt have patience for βlet me reinstall CUDA again.β
We tried:
- Running Ollama inside WSL (works, but no Unsloth optimizations)
- Docker-in-WSL with NVIDIA Container Toolkit (works, but adds a whole orchestration layer)
- Running Unsloth directly in WSL (dependency hell, version conflicts, crying)
- The ZimaBoard (remember that post? yeah, we remember too)
What we didnβt try β until we got desperate β was the obvious thing: put the model server on the Windows host, where the GPU is happy, and just talk to it from WSL.
The Revelation: Windows Host as AI Server
Hereβs the architecture that actually works:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Windows Host β
β βββββββββββββββββββ ββββββββββββββββββββββββββββ β
β β Unsloth Studio βββββββββΊβ netsh portproxy rule β β
β β (GPU accelerated) β 0.0.0.0:11434 β β β
β β localhost:11434 β localhost:11434 (relay) β β
β βββββββββββββββββββ ββββββββββββ¬ββββββββββββββββββ β
β β β
βββββββββββββββββββββββββββββββββββββββββββΌββββββββββββββββββββ
β WSL2 virtual switch
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β WSL 2 (Hermes) β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β Hermes Agent Framework β β
β β βββ Seek profile βββΊ HTTP request to β β
β β βββ Code profile βββΊ host.internal:11434 β β
β β βββ Scribe profile βββΊ (Unsloth Studio API) β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
The trick isnβt Docker. It isnβt a VPN. It isnβt some obscure WSL configuration. Itβs netsh interface portproxy β a built-in Windows utility thatβs been sitting there since Windows XP, quietly waiting for someone to need it.
The One Command That Makes It Work
On your Windows host (PowerShell as Administrator):
# Expose localhost:11434 (Unsloth Studio) to all interfaces so WSL can reach it
netsh interface portproxy add v4tov4 listenport=11434 listenaddress=0.0.0.0 connectport=11434 connectaddress=127.0.0.1
# Allow it through the Windows Firewall
netsh advfirewall firewall add rule name="Unsloth Studio WSL" dir=in action=allow protocol=tcp localport=11434
Now from inside WSL:
curl http://$(cat /etc/resolv.conf | grep nameserver | awk '{print $2}'):11434/api/tags
Boom. Your WSL instance is talking to Unsloth Studio on the Windows host. No Docker. No ssh tunnels. No VPN. Just two commands and a fundamental understanding that localhost in Windows and localhost in WSL are not the same place.
The IP you need is whateverβs in /etc/resolv.conf inside WSL β thatβs the WSL virtual switchβs gateway IP pointing back at the Windows host. Usually 172.x.x.x. You can set an alias in your shell profile so you never have to think about it again:
# ~/.bashrc or ~/.zshrc
export WINDOWS_HOST=$(cat /etc/resolv.conf | grep nameserver | awk '{print $2}')
What This Unlocks: Unsloth Studio on Windows
Why Unsloth Studio specifically? Because if youβre doing local inference on consumer hardware, Unsloth is the difference between βit worksβ and βit works fast.β
Unsloth implements hand-optimized kernels for model inference that squeeze 2-5x speedups out of standard Transformers. Itβs not a model β itβs an acceleration layer. When you run Qwen3 through Unsloth vs. stock Transformers, the difference is measurable in seconds per token, not milliseconds.
And hereβs the thing: Unsloth Studio runs flawlessly on Windows. Itβs a native Windows application with built-in model management, a chat interface, and an OpenAI-compatible local API. You download it, install it, point it at your models directory, and it justβ¦ works. Your GPU is fully utilized. Your VRAM is managed intelligently. You donβt fight with libcuda.so or LD_LIBRARY_PATH or any of the Linux GPU nonsense.
The API it exposes on :11434 is compatible with Ollamaβs format. That means anything that can talk to Ollama can talk to Unsloth Studio. Hermes. OpenWebUI. Continue.dev. Anything.

Model Comparison: What Actually Runs Here
When we got this working, the first thing we did was load up the models weβd been failing to run elsewhere. Hereβs the field report:
| Model | Size | VRAM | Speed (tok/s) | Notes |
|---|---|---|---|---|
| Gemma 4 9B (E2B) | 9B | ~6.5GB | 42 | The star. Handles Unslothβs E2B (extended context) variant with 256K context window. This is the model that froze our ZimaBoard. |
| Qwen3.5 | 0.6Bβ32B | 2GBβ20GB | 28β55 | Multiple variants. The 4B is the sweet spot for agent loops. |
| Llama 3.3 | 8B, 70B | 6GB, 40GB | 35, 8 | 8B runs warm. 70B needs a bigger card, but Unslothβs quantization makes it possible on 24GB. |
| DeepSeek R1 (distill) | 14B, 32B | 10GB, 20GB | 22, 12 | Reasoning models. Slower but smarter for complex tasks. |
| Phi-4 | 14B | 8GB | 30 | Microsoftβs latest. Surprisingly capable for its size. |
The Gemma4 E2B comparison that matters: On our previous attempts (Ollama-in-WSL, Ollama-on-ZimaBoard), Gemma4 either failed to load, hallucinated tool calls, or froze at 0% context. Through Unsloth Studio on Windows with direct GPU acceleration? 42 tokens per second, 256K context, zero tool hallucinations. The difference isnβt the model. Itβs the runtime.
The Architecture Note
Letβs talk about why this is the right architecture for a Windows+WSL dev setup, not just the convenient one.
Windows owns the hardware. Your GPU drivers, CUDA toolkit, and DirectML stack are designed for Windows first and Linux second. Every abstraction layer you add (WSL passthrough, Docker containers, remote boards) is a potential point of failure.
WSL owns the workflow. Your dev tools, your Hermes agents, your git repos, your Python environments β theyβre all happier in Linux. The filesystem is faster. The package manager works. The terminal doesnβt lie to you.
Portproxy bridges the gap without bridging the OS. Youβre not sharing CUDA between OSes. Youβre not fighting with WSLg or X11 forwarding. Youβre using the oldest networking primitive in the book: a TCP socket. Unsloth Studio speaks HTTP. Hermes speaks HTTP. The operating system they live on is irrelevant as long as packets can flow.
When we moved from βWSL runs everythingβ to βWindows runs GPU workloads, WSL runs agent logic,β our reliability went from ~60% (random CUDA errors, model load failures) to ~99%. The one percent is when Windows Update decides to reboot while youβre serving a model. Weβre working on that.
Practical Setup: From Zero to Agent Loop
If you want to replicate this exact setup:
1. Install Unsloth Studio on Windows
- Download from unsloth.ai
- Install, launch, download a model (Gemma 4 9B E2B is a great start)
- Verify the API is alive:
curl http://localhost:11434/api/tagsfrom PowerShell
2. Create the portproxy rule
# Run as Administrator
netsh interface portproxy add v4tov4 listenport=11434 listenaddress=0.0.0.0 connectport=11434 connectaddress=127.0.0.1
netsh advfirewall firewall add rule name="Unsloth Studio WSL" dir=in action=allow protocol=tcp localport=11434
3. Configure Hermes to point at it
In your Hermes provider config (or .env):
OLLAMA_BASE_URL=http://172.x.x.x:11434 # your WSL gateway IP
# Or dynamically:
OLLAMA_BASE_URL=http://$(cat /etc/resolv.conf | grep nameserver | awk '{print $2}'):11434
4. Test from Hermes
hermes chat --model gemma4:latest --message "Say hello from the Windows side"
5. Run a kanban task with a local model
hermes kanban work t_xxxx --model gemma4:latest

Benefits You Actually Feel
No more API bills. Once you own the inference, every agent run is free. We ran a full seek β code β scribe pipeline last night on Gemma4 and paid exactly $0.00.
No more context anxiety. 256K context windows mean your agent can read an entire codebase before answering. No more βplease summarize this in 4K tokens.β
No more hardware FOMO. Your existing GPU is probably enough. Unslothβs optimizations mean a 12GB card can comfortably run 32B models. A 24GB card can run 70B with quantization.
No more βwill it work today?β Windows drivers are stable. Unsloth Studio is a native app. The portproxy rules persist across reboots. This setup has the reliability of a kitchen appliance, not a research project.
Agent loops that actually loop. Remember when our ZimaBoard froze at 0% context for two minutes? That doesnβt happen here. Gemma4 under Unsloth handles the full Hermes XML tool schema without breaking a sweat.
What About the ZimaBoard?
Itβs still sitting there, warm and ready, running Ollama for casual chats. Itβs not dead β itβs just scoped. Single-shot inference, chat interfaces, coding copilot. Perfectly respectable work for a $300 box.
But for the heavy lifting β the multi-agent orchestration, the 256K context research tasks, the iterative coding workflows where an agent writes files and another agent reviews them β the Windows host + Unsloth Studio setup is the production stack.
The ZimaBoard is our development league pitcher. Unsloth Studio on Windows is our closer.
The Honest Verdict
| Approach | GPU Access | Context | Speed | Agent Loops | Cost |
|---|---|---|---|---|---|
| ZimaBoard + Ollama | 8GB RTX 3060 | 44K max | 4s/tok | β Freezes | $300 upfront |
| WSL + Ollama | CUDA passthrough | 128K | 12 tok/s | β οΈ Unreliable | $0 |
| Windows Host + Unsloth + portproxy | Full native | 256K | 42 tok/s | β Stable | $0 ongoing |
| Kimi K2.6 API | Their A100s | 128K | Instant | β Perfect | ~$0.001/tok |
For local agent work, the middle row wins. It gives you 90% of the cloud API experience at 0% of the per-token cost. The tradeoff is upfront hardware (which you probably already own) and a one-time setup (which you just read).
tl;dr for the Speedrunners
- Run Unsloth Studio on Windows. It owns the GPU and itβs happy there.
- Use
netsh portproxyto forward:11434to all interfaces. - Hit that port from WSL using the gateway IP in
/etc/resolv.conf. - Configure Hermesβ
OLLAMA_BASE_URLto point at it. - Run agent loops on Gemma4 E2B with 256K context for $0.
- Thank me later.
βThey said local multi-agent AI was impossible on consumer hardware. They just hadnβt tried letting Windows be Windows and Linux be Linux.β
β Cleetus π€‘
P.S. β If your Windows Firewall complains about allowing port 11434, tell it Cleetus said itβs fine. That usually works. If it doesnβt, add the rule properly. Firewalls are no joke.
P.P.S. β 42 tokens per second on a 4070 Ti. Thatβs not a benchmark number β thatβs what we measured while writing this post. The model was running, I was typing, and neither of us dropped a frame. Thatβs the Unsloth difference.
