Hosted image generators are easier, and for one-off images they are still the right answer. But open-weight models — the ones you download and run yourself — stopped being the compromise option. They now match hosted quality on most commercial work, and they do several things hosted tools cannot do at all.
This covers when the trade is worth making, and which models earn the disk space.
What Open-Weight Buys You
Consistency at volume. If you need 400 product shots in the same style, prompting a hosted tool 400 times gives you 400 variations. A fine-tuned local model gives you 400 versions of the same thing. That difference is the entire business case for most commercial use.
No per-image cost. Generating a thousand variations locally costs electricity. Generating a thousand variations on a hosted plan costs a thousand credits. Iteration-heavy workflows — the kind where you generate fifty images to pick one — get dramatically cheaper.
You can train on your own material. Brand assets, product photography, a specific illustration style. Hosted platforms increasingly restrict this. A local model has no terms of service standing between you and your training data.
Nothing leaves your network. Relevant if you work with unreleased products, confidential designs, or client material under an NDA.
The Models Worth Knowing
Qwen-Image
Alibaba's open-weight image model is the most interesting release of the year in this category. It is a 7B-parameter model, small enough to run on consumer hardware, and two things set it apart.
First, RGBA output — it can generate images with a real alpha channel, meaning transparent backgrounds come out of the model rather than being cut out afterwards. For anyone placing generated art onto product pages or into designs, that removes a whole step.
Second, multi-reference editing: you can give it several reference images and have it combine elements from them, rather than the single-image-reference workflow most editors offer.
It is free to use commercially, which matters more than it sounds. A lot of "open" models carry restrictions that quietly rule out commercial work.
FLUX
The quality benchmark for open-weight text-to-image in the earlier part of the year, and still excellent on typography and prompt adherence. Strong community fine-tune ecosystem, with variants tuned for realism, illustration and specific aesthetics. If you want the largest selection of pre-trained styles, this is where they live.
Stable Diffusion
Still the most extensible by a wide margin. LoRAs, ControlNet, IP-Adapter, regional prompting — the entire tooling ecosystem was built here first. If your workflow needs precise structural control (pose, depth, layout), the plugins exist here and nowhere else in the same depth.
ComfyUI
Not a model, but the reason any of this is usable. It is a node-based interface for building image pipelines visually — load a model, add a ControlNet, route a LoRA, batch it, save. Anything you build can be saved and re-run, which turns a one-off trick into a repeatable production step.
What Fine-Tuning Actually Gives You
Two levels, and they are very different amounts of work:
LoRA (low-rank adaptation) is the practical one. You supply 15–40 reference images, train for under an hour on a single consumer GPU, and get a small file that teaches the model a subject or style. This is how people put a specific person, product or art style into a model. It is achievable in an afternoon.
Full fine-tuning retrains the model itself. It needs substantially more data, more compute and more patience, and it is worth it only if you are building a product with a genuinely distinctive visual identity.
The honest caveat: fine-tuning solves consistency, not quality. A LoRA trained on mediocre images produces consistent mediocrity.
When to Stay Hosted
Do not self-host if any of these is true:
- You need a handful of images a month. Midjourney will produce better results in less time than you will spend configuring anything.
- You have no suitable hardware. An 8 GB GPU is workable; 4 GB is painful; CPU-only generation is a test of patience.
- You want someone else to handle updates. Local models do not improve themselves. New releases are yours to find and install.
Browse the full selection in our Image category.
Frequently Asked Questions
Can I use open-weight image models for commercial work?
Usually yes, but check the specific licence — this is the one place people get burned. Models labelled open-weight vary: some permit commercial use freely, others restrict it above a revenue threshold or require attribution. Qwen-Image permits commercial use; always verify for the specific model you plan to ship with.
What hardware do I need?
A GPU with 8 GB of VRAM handles most open-weight image generation comfortably. 12 GB or more opens up larger models and higher resolutions. Apple Silicon with 16 GB+ unified memory works, though slower than a comparable discrete GPU.
Is quality really comparable to Midjourney?
On photorealism and stylised illustration, the gap has closed to the point where most people cannot identify which produced a given image. Midjourney still leads on aesthetic defaults — its output looks good with less prompting effort, which is a real advantage if you are not going to iterate.
What is a LoRA, in plain terms?
A small add-on file that teaches a model a specific subject or style without retraining the whole thing. You train it on 15–40 of your own images in under an hour, then load it alongside the base model. It is the most practical way to get consistent, on-brand output.
Do I need ComfyUI?
Not for simple generation — most workflows run fine in a basic interface. You want ComfyUI once your process has steps: reference images, structural control, upscaling, batching. It turns a sequence of manual actions into one pipeline you can re-run.
