By Tamsi Besson ·
Why I run a remote Qwen 3.6 27B server instead of paying Cursor
Free tokens instead of Cursor models — Qwen 3.6 27B on AWS via vLLM (4-bit), Ollama locally for git-mentor.
- LLM
- Qwen
- AWS
- vLLM
- Cursor
Cursor is a great agentic IDE. Its native models, though, got frustrating fast: quotas, per-token billing, agent sessions stopping at the worst time, and a quality/price ratio that’s hard to defend when you chain code reviews, refactors, and MCP experiments all day.
The actual problem
It wasn’t “I want to self-host an LLM for fun.” It was: I want to use the agent without checking the token meter every ten minutes. My projects (MCP, Gradio prototypes, large GitHub diffs) eat a lot of context. Paying Cursor for every iteration is fine occasionally — not on a loop for side projects and open source.
What the remote server gives me
- Free inference tokens: I pay for GPU hours, not per prompt. I can rerun the agent ten times on the same bug without guilt.
- Qwen 3.6 27B on an OpenAI-compatible API — Cursor, my MCP servers, and scripts share one endpoint.
- Same IDE workflow; only the model changes. No need to leave Cursor.
- Independence: if pricing or limits shift again, my stack stays mine.
How it’s set up (briefly)
GPU instance on AWS (48 GB VRAM), vLLM serving Qwen 3.6 27B, TLS reverse proxy + Bearer token in front. vLLM exposes the OpenAI-compatible API (/v1/chat/completions) — Cursor, ai-code-reviewer-mcp, and redbee-mcp all point at the same endpoint with shared OPENAI_API_BASE / OPENAI_API_KEY.
Why vLLM on AWS (not Ollama here)
For a 27B in agent production — long context, parallel requests, uptime — vLLM fits better: throughput, batching, stable API. Ollama shines elsewhere (see git-mentor below). On AWS I load the quantized model once; I pay for the GPU instance, not Cursor tokens.
Why 4-bit (AWQ / GPTQ)
A 27B in FP16 needs ~54 GB VRAM — out of reach on a single “normal” card. With 4-bit quantization (AWQ or GPTQ, formats vLLM loads natively), weights land around 16 GB: the model fits on GPU with headroom for long context and several agent requests. You lose a bit of nuance vs full precision, but for code review, refactors, and MCP calls the gap is rarely blocking — especially compared to a large cloud model billed per token.
- FP16: best quality, 3× VRAM — makes sense if you have 80 GB+ or multiple GPUs.
- AWQ / GPTQ 4-bit: solid quality/price/latency tradeoff for daily vLLM use.
- Q8: middle ground if you want a bit more quality without doubling the EC2 bill.
# On the AWS instance (vLLM + AWQ weights example)
vllm serve Qwen/Qwen3-27B-Instruct-AWQ \
--quantization awq \
--host 127.0.0.1 --port 8000
# Cursor / MCP (via TLS proxy)
OPENAI_API_BASE=https://llm.example.com/v1
OPENAI_API_KEY=<proxy-token>
OPENAI_MODEL=Qwen/Qwen3-27B-Instruct-AWQLocal Ollama for git-mentor
git-mentor is a different use case: GitHub profile analysis with a PAT that shouldn’t hit a remote server if I can avoid it, short sessions, a smaller model is enough. There I use Ollama on my machine — zero cloud cost, works offline, localhost:11434/v1. The big Qwen on AWS stays for Cursor and MCP; git-mentor stays local-first by design.
# Mac / laptop — git-mentor locally
ollama run llama3.2
export OPENAI_API_BASE=http://localhost:11434/v1
export OPENAI_API_KEY=ollama
git-mentor analyze --user TamsiBottom line
Two complementary stacks: vLLM on AWS for the “full throttle” agent (free inference tokens), Ollama locally for git-mentor. AWS setup took an afternoon; the real win is mental and economic. Cursor stays my interface; the heavy brain runs on my GPU in the cloud, the light one on the laptop. Follow-up: Qwen 3.8 27B — /blog/qwen-3-8-27b.