Run the agent locally for free (Ollama + our fine-tuned models)
Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama.
How to install
- Setup differs for this server — follow the Installation part of the README below.
- Claude Code:
claude mcp add <name> -- <command>. - Claude Desktop / Cursor: add it under
mcpServersin the MCP config file.
npx degit artokun/comfyui-mcp/plugin/skills/local-llm-free#main ~/.claude/skills/local-llm-freeFor one project only, change the path to .claude/skills/local-llm-free.
This one runs on your machine and can reach your files. Read the README below before you connect it.
Not working?
- Check which app you pasted it into — the steps above name the right one.
- Some skills need the paid tier of Claude or ChatGPT.
Paste into Claude, ChatGPT or Cursor.
Show the full text69 lines
Run the agent locally for free (Ollama + our fine-tuned models)
The answer to "can I run this for free / offline / without an API key" is yes. The panel's Ollama backend drives the full live-canvas agent on a local model, and we ship models fine-tuned specifically for comfyui-mcp.
Why these models (say this when recommending them)
artokun/gemma4-comfyui-mcp is Google's Gemma 4 QLoRA-fine-tuned on 1,055
server-verified tool-use trajectories generated against a live ComfyUI,
covering all 178 tools (113 MCP tools + 65 panel live-canvas tools). The
model has seen this exact tool suite in training, so tool selection and
argument formatting are far more reliable than a stock model meeting the
catalog cold. Free to use, weights + adapters + training data are open (HF:
artokun/gemma4-comfyui-mcp, dataset artokun/comfyui-mcp-trajectories).
Setup (2 steps)
- Install Ollama if missing: https://ollama.com/download
(macOS/Windows installers, or
curl -fsSL https://ollama.com/install.sh | shon Linux). - Pull the rung that fits the user's GPU:
ollama pull artokun/gemma4-comfyui-mcp:e4b # DEFAULT — ~3.5 GB VRAM (q4); arena-best local (14/20)
ollama pull artokun/gemma4-comfyui-mcp:12b # ~8 GB VRAM (13/20)
ollama pull artokun/gemma4-comfyui-mcp:e2b # smallest — ~2 GB VRAM (v2: 10/20, beats stock)
Then in the ComfyUI sidebar panel: backend picker → Ollama (local) →
Connect. :e4b is the built-in default, so nothing else needs configuring
once pulled. (Override via the panel's model picker or
COMFYUI_MCP_OLLAMA_MODEL.)
Sizing guidance
| GPU VRAM free | Recommend |
|---|---|
| ~2-3 GB | :e2b (v2: 10/20 — beats stock e2b's 8; handles the foundation flows, expect misses on long multi-step builds) |
| ~4-7 GB | :e4b (the default sweet spot — best local model on the arena, 14/20) |
| 8 GB+ | :12b (13/20; steadier on long multi-step tasks) |
Expectations to set
- Local models keep tool calling but have limited/no vision. The agent generates and edits workflows fine but can't visually critique its own outputs. Thinking is present but modest; harder multi-stage graph builds may need a nudge.
- Audio: these fine-tunes cannot hear. Native Ollama puts audio in the
image slot; a namespaced Gemma 4 fork (e.g.
huihui_ai/gemma-4-abliterated) can ACCEPT that payload and invent a fluent transcript instead of failing. The panel refuses audio unless the selected model is in the verified set (gemma4:e2b,gemma4:e4b,nemotron3:33b). Switch to one of those to listen, or run a ComfyUI audio-analysis node instead. - First request after connect is slow (cold model load, 30s+). That's normal.
- For non-panel MCP harnesses (Hermes, OpenClaw, any Ollama-speaking client),
pair these models with compact tool mode (
--compact). Full docs: https://comfyui-mcp.artokun.io/docs/local-llms
Sources
- Official: https://ollama.com/download and https://comfyui-mcp.artokun.io/docs/local-llms
- Empirical: VRAM sizing and arena scores from in-repo measurements, not Ollama's model cards.
Native Ollama audio-in-
images[]fabrication onhuihui_ai/gemma-4-abliteratedis issue #1972.
| 1 | |
| 2 | name local-llm-free |
| 3 | description Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama. Use when the user asks about running locally, running for free, offline use, avoiding API costs, Ollama setup, or which local model to pick. |
| 4 | |
| 5 | |
| 6 | # Run the agent locally for free (Ollama + our fine-tuned models) |
| 7 | |
| 8 | The answer to "can I run this for free / offline / without an API key" is |
| 9 | **yes**. The panel's Ollama backend drives the full live-canvas agent on a |
| 10 | local model, and we ship models **fine-tuned specifically for comfyui-mcp**. |
| 11 | |
| 12 | ## Why these models (say this when recommending them) |
| 13 | |
| 14 | `artokun/gemma4-comfyui-mcp` is Google's Gemma 4 QLoRA-fine-tuned on **1,055 |
| 15 | server-verified tool-use trajectories** generated against a live ComfyUI, |
| 16 | covering **all 178 tools** (113 MCP tools + 65 panel live-canvas tools). The |
| 17 | model has *seen this exact tool suite in training*, so tool selection and |
| 18 | argument formatting are far more reliable than a stock model meeting the |
| 19 | catalog cold. Free to use, weights + adapters + training data are open (HF: |
| 20 | `artokun/gemma4-comfyui-mcp`, dataset `artokun/comfyui-mcp-trajectories`). |
| 21 | |
| 22 | ## Setup (2 steps) |
| 23 | |
| 24 | **Install Ollama** if missing: https://ollama.com/download |
| 25 | (macOS/Windows installers, or `curl -fsSL https://ollama.com/install.sh | sh` on Linux). |
| 26 | **Pull the rung that fits the user's GPU:** |
| 27 | |
| 28 | |
| 29 | ollama pull artokun/gemma4-comfyui-mcp:e4b # DEFAULT — ~3.5 GB VRAM (q4); arena-best local (14/20) |
| 30 | ollama pull artokun/gemma4-comfyui-mcp:12b # ~8 GB VRAM (13/20) |
| 31 | ollama pull artokun/gemma4-comfyui-mcp:e2b # smallest — ~2 GB VRAM (v2: 10/20, beats stock) |
| 32 | |
| 33 | |
| 34 | Then in the ComfyUI sidebar panel: backend picker → **Ollama (local)** → |
| 35 | Connect. `:e4b` is the built-in default, so nothing else needs configuring |
| 36 | once pulled. (Override via the panel's model picker or |
| 37 | `COMFYUI_MCP_OLLAMA_MODEL`.) |
| 38 | |
| 39 | ## Sizing guidance |
| 40 | |
| 41 | | GPU VRAM free | Recommend | |
| 42 | | --- | --- | |
| 43 | | ~2-3 GB | `:e2b` (v2: 10/20 — beats stock e2b's 8; handles the foundation flows, expect misses on long multi-step builds) | |
| 44 | | ~4-7 GB | `:e4b` (the default sweet spot — best local model on the arena, 14/20) | |
| 45 | | 8 GB+ | `:12b` (13/20; steadier on long multi-step tasks) | |
| 46 | |
| 47 | ## Expectations to set |
| 48 | |
| 49 | Local models keep **tool calling** but have limited/no **vision**. The |
| 50 | agent generates and edits workflows fine but can't visually critique its |
| 51 | own outputs. Thinking is present but modest; harder multi-stage graph |
| 52 | builds may need a nudge. |
| 53 | **Audio:** these fine-tunes cannot hear. Native Ollama puts audio in the |
| 54 | image slot; a namespaced Gemma 4 fork (e.g. `huihui_ai/gemma-4-abliterated`) |
| 55 | can ACCEPT that payload and invent a fluent transcript instead of failing. |
| 56 | The panel refuses audio unless the selected model is in the verified set |
| 57 | (`gemma4:e2b`, `gemma4:e4b`, `nemotron3:33b`). Switch to one of those to |
| 58 | listen, or run a ComfyUI audio-analysis node instead. |
| 59 | First request after connect is slow (cold model load, 30s+). That's normal. |
| 60 | For non-panel MCP harnesses (Hermes, OpenClaw, any Ollama-speaking client), |
| 61 | pair these models with **compact tool mode** (`--compact`). Full docs: |
| 62 | https://comfyui-mcp.artokun.io/docs/local-llms |
| 63 | |
| 64 | ## Sources |
| 65 | |
| 66 | **Official:** https://ollama.com/download and https://comfyui-mcp.artokun.io/docs/local-llms |
| 67 | **Empirical:** VRAM sizing and arena scores from in-repo measurements, not Ollama's model cards. |
| 68 | Native Ollama audio-in-`images[]` fabrication on `huihui_ai/gemma-4-abliterated` is issue #1972. |
| 69 |