Run the agent locally for free (Ollama + our fine-tuned models)

Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama.

How to install

How to install

  1. Setup differs for this server — follow the Installation part of the README below.
  2. Claude Code: claude mcp add <name> -- <command>.
  3. Claude Desktop / Cursor: add it under mcpServers in the MCP config file.
Claude Code — installs the whole folder, not just SKILL.md
npx degit artokun/comfyui-mcp/plugin/skills/local-llm-free#main ~/.claude/skills/local-llm-free

For one project only, change the path to .claude/skills/local-llm-free.

This one runs on your machine and can reach your files. Read the README below before you connect it.

Not working?
  • Check which app you pasted it into — the steps above name the right one.
  • Some skills need the paid tier of Claude or ChatGPT.
Step-by-step guide with screenshots · Ask in the forum

Paste into Claude, ChatGPT or Cursor.

Show the full text69 lines
local-llm-free/SKILL.md69 lines3.5 KBpushed 27d agoRawView on GitHub

Run the agent locally for free (Ollama + our fine-tuned models)

The answer to "can I run this for free / offline / without an API key" is yes. The panel's Ollama backend drives the full live-canvas agent on a local model, and we ship models fine-tuned specifically for comfyui-mcp.

Why these models (say this when recommending them)

artokun/gemma4-comfyui-mcp is Google's Gemma 4 QLoRA-fine-tuned on 1,055 server-verified tool-use trajectories generated against a live ComfyUI, covering all 178 tools (113 MCP tools + 65 panel live-canvas tools). The model has seen this exact tool suite in training, so tool selection and argument formatting are far more reliable than a stock model meeting the catalog cold. Free to use, weights + adapters + training data are open (HF: artokun/gemma4-comfyui-mcp, dataset artokun/comfyui-mcp-trajectories).

Setup (2 steps)

  1. Install Ollama if missing: https://ollama.com/download (macOS/Windows installers, or curl -fsSL https://ollama.com/install.sh | sh on Linux).
  2. Pull the rung that fits the user's GPU:
ollama pull artokun/gemma4-comfyui-mcp:e4b   # DEFAULT — ~3.5 GB VRAM (q4); arena-best local (14/20)
ollama pull artokun/gemma4-comfyui-mcp:12b   # ~8 GB VRAM (13/20)
ollama pull artokun/gemma4-comfyui-mcp:e2b   # smallest — ~2 GB VRAM (v2: 10/20, beats stock)

Then in the ComfyUI sidebar panel: backend picker → Ollama (local) → Connect. :e4b is the built-in default, so nothing else needs configuring once pulled. (Override via the panel's model picker or COMFYUI_MCP_OLLAMA_MODEL.)

Sizing guidance

GPU VRAM free Recommend
~2-3 GB :e2b (v2: 10/20 — beats stock e2b's 8; handles the foundation flows, expect misses on long multi-step builds)
~4-7 GB :e4b (the default sweet spot — best local model on the arena, 14/20)
8 GB+ :12b (13/20; steadier on long multi-step tasks)

Expectations to set

  • Local models keep tool calling but have limited/no vision. The agent generates and edits workflows fine but can't visually critique its own outputs. Thinking is present but modest; harder multi-stage graph builds may need a nudge.
  • Audio: these fine-tunes cannot hear. Native Ollama puts audio in the image slot; a namespaced Gemma 4 fork (e.g. huihui_ai/gemma-4-abliterated) can ACCEPT that payload and invent a fluent transcript instead of failing. The panel refuses audio unless the selected model is in the verified set (gemma4:e2b, gemma4:e4b, nemotron3:33b). Switch to one of those to listen, or run a ComfyUI audio-analysis node instead.
  • First request after connect is slow (cold model load, 30s+). That's normal.
  • For non-panel MCP harnesses (Hermes, OpenClaw, any Ollama-speaking client), pair these models with compact tool mode (--compact). Full docs: https://comfyui-mcp.artokun.io/docs/local-llms

Sources

1---
2name: local-llm-free
3description: Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama. Use when the user asks about running locally, running for free, offline use, avoiding API costs, Ollama setup, or which local model to pick.
4---
5 
6# Run the agent locally for free (Ollama + our fine-tuned models)
7 
8The answer to "can I run this for free / offline / without an API key" is
9**yes**. The panel's Ollama backend drives the full live-canvas agent on a
10local model, and we ship models **fine-tuned specifically for comfyui-mcp**.
11 
12## Why these models (say this when recommending them)
13 
14`artokun/gemma4-comfyui-mcp` is Google's Gemma 4 QLoRA-fine-tuned on **1,055
15server-verified tool-use trajectories** generated against a live ComfyUI,
16covering **all 178 tools** (113 MCP tools + 65 panel live-canvas tools). The
17model has *seen this exact tool suite in training*, so tool selection and
18argument formatting are far more reliable than a stock model meeting the
19catalog cold. Free to use, weights + adapters + training data are open (HF:
20`artokun/gemma4-comfyui-mcp`, dataset `artokun/comfyui-mcp-trajectories`).
21 
22## Setup (2 steps)
23 
241. **Install Ollama** if missing: https://ollama.com/download
25 (macOS/Windows installers, or `curl -fsSL https://ollama.com/install.sh | sh` on Linux).
262. **Pull the rung that fits the user's GPU:**
27 
28```bash
29ollama pull artokun/gemma4-comfyui-mcp:e4b # DEFAULT — ~3.5 GB VRAM (q4); arena-best local (14/20)
30ollama pull artokun/gemma4-comfyui-mcp:12b # ~8 GB VRAM (13/20)
31ollama pull artokun/gemma4-comfyui-mcp:e2b # smallest — ~2 GB VRAM (v2: 10/20, beats stock)
32```
33 
34Then in the ComfyUI sidebar panel: backend picker → **Ollama (local)**
35Connect. `:e4b` is the built-in default, so nothing else needs configuring
36once pulled. (Override via the panel's model picker or
37`COMFYUI_MCP_OLLAMA_MODEL`.)
38 
39## Sizing guidance
40 
41| GPU VRAM free | Recommend |
42| --- | --- |
43| ~2-3 GB | `:e2b` (v2: 10/20 — beats stock e2b's 8; handles the foundation flows, expect misses on long multi-step builds) |
44| ~4-7 GB | `:e4b` (the default sweet spot — best local model on the arena, 14/20) |
45| 8 GB+ | `:12b` (13/20; steadier on long multi-step tasks) |
46 
47## Expectations to set
48 
49- Local models keep **tool calling** but have limited/no **vision**. The
50 agent generates and edits workflows fine but can't visually critique its
51 own outputs. Thinking is present but modest; harder multi-stage graph
52 builds may need a nudge.
53- **Audio:** these fine-tunes cannot hear. Native Ollama puts audio in the
54 image slot; a namespaced Gemma 4 fork (e.g. `huihui_ai/gemma-4-abliterated`)
55 can ACCEPT that payload and invent a fluent transcript instead of failing.
56 The panel refuses audio unless the selected model is in the verified set
57 (`gemma4:e2b`, `gemma4:e4b`, `nemotron3:33b`). Switch to one of those to
58 listen, or run a ComfyUI audio-analysis node instead.
59- First request after connect is slow (cold model load, 30s+). That's normal.
60- For non-panel MCP harnesses (Hermes, OpenClaw, any Ollama-speaking client),
61 pair these models with **compact tool mode** (`--compact`). Full docs:
62 https://comfyui-mcp.artokun.io/docs/local-llms
63 
64## Sources
65 
66- **Official:** https://ollama.com/download and https://comfyui-mcp.artokun.io/docs/local-llms
67- **Empirical:** VRAM sizing and arena scores from in-repo measurements, not Ollama's model cards.
68 Native Ollama audio-in-`images[]` fabrication on `huihui_ai/gemma-4-abliterated` is issue #1972.
69 

Discussion

Alternatives

Also in Illustration & art