Stable Diffusion WebUI is still one of the most searched local AI tools on the internet, and the ecosystem around it has split into four distinct paths in 2026: the original AUTOMATIC1111 repository, the once-dominant Forge fork, the community-driven reForge project, and Forge Classic, which is now the most actively maintained branch of the family. Picking the wrong one wastes an afternoon chasing dependency errors, and plenty of guides floating around the web still point readers at a repository that stopped shipping updates more than a year ago. This tutorial walks through the entire setup from a clean machine to a working local image generator with ControlNet, LoRA support, and a scriptable API, using the branch that is actually still receiving updates as of October 2026.
By the end you will have a fully functional installation, a batch-generation Python script that talks to the WebUI over its REST API, and a troubleshooting reference for the errors that trip up almost everyone on their first install. Total setup time runs 45 to 90 minutes depending on your download speed and whether your GPU driver is already current. None of this requires a subscription, a credit card, or an internet connection once the checkpoints are downloaded, which is the main reason people still bother with a local install when hosted APIs are a single curl command away.
This guide uses thirteen concrete steps, grouped into logical stages so you are not jumping between unrelated settings mid-install. Each stage builds on the last: environment setup first, then the base install, then performance tuning, then the extensions and automation layer that turn a basic image generator into something you can actually build a workflow around.
Why Stable Diffusion WebUI Still Matters in 2026
Cloud AI image APIs like Nano Banana Pro, GPT Image 2.5, and FLUX.2 get most of the headlines because they require zero setup and scale instantly, and the site keeps a running comparison of where each one stands in the best AI models roundup. But a local stable diffusion webui install still wins on three fronts: no per-image cost after the hardware is paid for, no content filters blocking legitimate commercial or creative work, and full control over which checkpoint, LoRA, or ControlNet model actually runs the generation. For anyone doing high-volume product photography, game asset iteration, or repeated batch renders, the economics flip in favor of local compute within a few thousand images.
The original AUTOMATIC1111/stable-diffusion-webui repository on GitHub sits at 165,170 stars as of this writing, with 2,512 open issues and a last code push dated March 2, 2026, which puts it firmly in maintenance mode rather than active development. By contrast, Comfy-Org/ComfyUI has climbed to 135,758 stars with commits landing the same day this article was published, which is why node-based workflows have become the default recommendation for anyone chasing the newest model architectures like FLUX or Qwen-Image. For people who specifically want the classic AUTOMATIC1111-style tabbed interface rather than a node graph, Forge Classic is the build worth installing, and that is the path this guide follows.
It helps to understand why the fork even exists. AUTOMATIC1111’s original codebase loads the entire model pipeline into VRAM up front, which made SDXL painful on anything under 10GB. Forge’s contribution, carried forward into Forge Classic, was a dynamic memory-management system that swaps model components between VRAM and system RAM on demand, which is the reason SDXL became usable on 6-8GB consumer cards in the first place. That architectural change is also why Forge-family builds and plain AUTOMATIC1111 are not perfectly interchangeable: some older extensions written specifically for AUTOMATIC1111’s memory model do not load cleanly on Forge Classic without an update from their maintainer.
AUTOMATIC1111 vs Forge vs Forge Classic vs reForge: Picking Your Build
Before touching a terminal, it helps to know which repository you are actually installing, because the name “Forge” alone refers to at least three different projects with different maintenance states. The original Forge repository by lllyasviel, which pioneered the memory-management rewrite that let SDXL run on 4GB cards, has not had a meaningful commit since July 31, 2025. That makes it over a year stale, and its own community now points newcomers toward its successors.
reForge, maintained by the developer Panchovix, keeps pulling upstream AUTOMATIC1111 changes and newer samplers into the Forge architecture, and its repository shows activity as recently as April 14, 2026. Forge Classic, maintained separately by Haoming02, had a commit land on October 1, 2026 — the same day this guide was written — and is explicitly positioned as the actively maintained continuation of the Forge codebase for SD 1.5 and SDXL checkpoints. The table below lays out the comparison with the exact numbers pulled directly from each project’s GitHub repository.
One detail worth flagging before you commit to a fork: Forge Classic itself maintains two branches with very different targets. The “classic” branch, used throughout this tutorial, deliberately limits scope to SD 1.5 and SDXL checkpoints and only accepts critical bug fixes going forward, which makes it predictable and low-risk for a production setup. A separate “neo” branch under active development targets Python 3.13 and a broader set of newer architectures, but it moves faster and is more likely to break between updates. Unless you specifically need neo’s newer feature set, classic is the safer starting point.
| Project | GitHub stars | Last commit | Status | Best for |
|---|---|---|---|---|
| AUTOMATIC1111/stable-diffusion-webui | 165,170 | Mar 2, 2026 | Maintenance mode | Maximum extension compatibility |
| lllyasviel/stable-diffusion-webui-forge | 13,043 | Jul 31, 2025 | Largely inactive | Legacy installs only |
| Panchovix/stable-diffusion-webui-reForge | 1,042 | Apr 14, 2026 | Actively maintained | Newer samplers, SDXL/SD1.5 |
| Haoming02/sd-webui-forge-classic | 1,773 | Oct 1, 2026 | Actively maintained | This tutorial’s recommended build |
| Comfy-Org/ComfyUI | 135,758 | Oct 1, 2026 | Very active | FLUX, Qwen-Image, node workflows |
If your goal is the familiar AUTOMATIC1111 tab layout with ongoing bug fixes and current PyTorch support, install Forge Classic’s “classic” branch, which is what every step below targets. If you specifically need FLUX, Qwen-Image, or another transformer-based model with the widest possible node support, read the site’s existing ComfyUI tutorial for SDXL and FLUX workflows instead, since ComfyUI’s ecosystem is built specifically around those architectures. The site also has a dedicated walkthrough for running FLUX locally in ComfyUI and one for self-hosting Qwen-Image-2.1, both of which cover architectures that Forge Classic does not support as cleanly.
Prerequisites and System Requirements
Stable Diffusion WebUI Forge Classic’s classic branch recommends Python 3.11.9 specifically, not the newest Python release. Installing Python 3.12 or 3.13 for the classic branch can trigger package build failures during first launch, so stick to the documented version even if a newer interpreter is already on your system. You will also need Git, roughly 25 to 30GB of free SSD space for the repository plus one checkpoint, and an NVIDIA GPU for any reasonable generation speed, though CPU-only and AMD-via-ROCm paths exist with reduced performance.
| Checkpoint type | Minimum VRAM | Recommended VRAM | Typical disk size |
|---|---|---|---|
| Stable Diffusion 1.5 | 4GB | 6-8GB | 2-4GB per checkpoint |
| SDXL base/refiner | 6-8GB | 12GB | 6-7GB per checkpoint |
| Stable Diffusion 3.5 | 10-16GB | 16-24GB | 10-17GB per checkpoint |
| FLUX.1 (quantized) | 12-16GB | 24GB+ | 12-23GB per checkpoint |
- Python 3.11.9 (classic branch) — download from python.org
- Git for Windows, macOS, or Linux
- NVIDIA GPU with current driver, or Apple Silicon / AMD ROCm for non-NVIDIA paths
- At least 16GB system RAM (32GB avoids disk paging on larger batches)
- 25-30GB free SSD space before downloading any checkpoints
- Optional: a free Hugging Face or Civitai account for checkpoint downloads
Step 1-2: Install Python, Git, and Your GPU Driver
Start by installing Python 3.11.9 from python.org and checking “Add Python to PATH” during the Windows installer, since the launch scripts call python directly from the terminal. On macOS or Linux, use a version manager like pyenv to pin 3.11 without disturbing your system Python. Next, install Git, which is how you will pull the Forge Classic repository and later update it.
Confirm your NVIDIA driver is current enough to support the CUDA version Forge Classic’s classic branch ships with (PyTorch 2.10.0 built against CUDA 12.9). Download the latest driver and CUDA toolkit directly from NVIDIA’s CUDA downloads page if nvidia-smi reports an older version below. An outdated driver is one of the most common causes of a silent crash on first launch, so it is worth checking before going further rather than debugging it after a failed install.
python --version
git --version
nvidia-smi
The nvidia-smi command should print your GPU model, driver version, and the maximum CUDA version your driver supports. If that number is below 12.x, update your driver from NVIDIA’s site before continuing, or the installer will fail partway through dependency resolution.
Step 3-4: Clone the Repository and Build Your Virtual Environment
Clone the classic branch specifically. The repository hosts multiple branches, including an experimental “neo” branch aimed at newer Python and broader model support, but classic is the one that focuses on SD 1.5 and SDXL stability and only receives critical bug fixes, which makes it the safer choice for a first install.
git clone https://github.com/Haoming02/sd-webui-forge-classic --branch classic
cd sd-webui-forge-classic
Forge Classic supports the uv package manager as a much faster alternative to pip for the dependency install. If you have uv installed already, create the virtual environment with it; otherwise the default launch script will fall back to building a standard venv with pip, which works fine but takes noticeably longer on first boot.
# Recommended: faster install with uv
uv venv venv --python 3.11 --seed
# Then add --uv to the COMMANDLINE_ARGS line in webui-user.bat
# to make the installer use uv pip instead of regular pip
Step 5-6: First Boot and Downloading Your First Checkpoint
Launch the WebUI for the first time by running webui-user.bat on Windows, or webui.sh on macOS/Linux. The first launch automatically downloads the remaining Python dependencies and can take 10 to 20 minutes depending on your connection, since it is pulling PyTorch with CUDA support, which alone runs several gigabytes.
While that installs, grab your first checkpoint. SD 1.5-based checkpoints remain the lightest and fastest way to confirm the install works, while SDXL checkpoints give noticeably sharper output at the cost of more VRAM and slower generation. If you want to compare against the newer Stable Diffusion 3.5 architecture once your base install is working, the site’s Stable Diffusion 3.5 guide covers the differences in detail. Place the downloaded .safetensors file in the correct models folder before your first generation attempt.
Two sources cover almost every checkpoint you will need. Stability AI’s own Hugging Face page hosts the official base SD 1.5, SDXL, and SD 3.5 weights with no sign-in required for most files, and it is the safest first stop since the models come directly from the company that trains them. Civitai hosts the much larger long tail of community fine-tuned checkpoints and LoRAs built on top of those bases, which is where most stylistic variety lives, though quality and licensing vary file by file, so check the model card before using anything commercially.
# Checkpoint files go here, relative to the repo root:
sd-webui-forge-classic/models/Stable-diffusion/your-checkpoint.safetensors
Once the terminal prints a local URL such as http://127.0.0.1:7860, open it in your browser. You should see the familiar tabbed interface with txt2img, img2img, Extras, and Settings tabs along the top, confirming the install completed successfully. If the page loads but the checkpoint dropdown in the top-left corner is empty, click the small refresh icon next to it; Forge Classic only scans the models folder at launch and after a manual refresh, not continuously.
Step 7: VRAM and Performance Command-Line Flags
Before generating anything, open webui-user.bat in a text editor and set your COMMANDLINE_ARGS line based on your GPU’s VRAM. These flags trade generation speed for lower memory usage, and getting them wrong is the single most common reason people report crashes or out-of-memory errors on their first few generations.
| Flag | Effect | When to use it |
|---|---|---|
| –xformers | Memory-efficient attention, faster generation | Most NVIDIA GPUs except RTX 50-series, where it is unsupported |
| –medvram | Splits model across VRAM/RAM, moderate slowdown | 6-8GB VRAM cards running SDXL |
| –lowvram | Aggressive memory splitting, significant slowdown | 4GB or less VRAM |
| –uv | Uses uv instead of pip for dependency installs | Every install, for faster setup and updates |
| –uv-symlink | Symlinks packages instead of copying them | Cuts install footprint from roughly 7GB to about 100MB |
| –sage | Installs SageAttention for faster sampling | RTX 40-series and newer, when supported |
A typical 8GB-VRAM line looks like set COMMANDLINE_ARGS=--medvram --xformers --uv. Save the file, close any running WebUI instance, and relaunch for the flags to take effect.
Step 8: Generate Your First Image With txt2img
With the checkpoint loaded in the top-left dropdown, switch to the txt2img tab. Enter a descriptive prompt, leave the sampler on its default (DPM++ 2M Karras is a reliable starting point for both SD 1.5 and SDXL), set steps between 20 and 30, and click Generate. A first image on an RTX 4070-class GPU with an SD 1.5 checkpoint typically completes in 3 to 6 seconds at 512×512; SDXL at 1024×1024 runs closer to 8 to 15 seconds on the same hardware.
Example output you should expect on a clean install: a 512×512 or 1024×1024 PNG saved automatically to the outputs/txt2img-images folder, named with a timestamp, alongside a matching .txt file recording the exact prompt, seed, sampler, and CFG scale used, which is what makes results reproducible later.
Understanding Samplers, CFG Scale, and Steps
The three settings that affect output quality the most are the ones beginners tend to leave untouched, which is unfortunate since getting them wrong is the most common reason a first-time user dismisses the whole tool as low quality. Steps control how many denoising passes the model runs; below 15 steps, images often look muddy or unfinished, while pushing past 40-50 steps rarely improves quality further and just slows generation down. The 20-30 range is the sweet spot for almost every modern sampler.
CFG scale (classifier-free guidance) controls how strictly the model follows your prompt versus how much creative freedom it takes. A low CFG scale around 3-5 produces looser, sometimes more natural-looking results; a high CFG scale above 12 forces stricter prompt adherence but can introduce oversaturated colors and artifacts. Most checkpoints are tuned to look best between 5 and 8.
The sampler itself determines the mathematical path the model takes from random noise to a finished image. DPM++ 2M Karras remains the most broadly reliable default across SD 1.5 and SDXL checkpoints. Euler a produces slightly more varied, looser results at the same step count and is worth trying for creative work where exact reproducibility matters less. DPM++ SDE Karras tends to produce sharper fine detail but runs somewhat slower per step. None of these choices are wrong; they are tradeoffs, and the fastest way to find your preference is running the same prompt and seed across two or three samplers using the X/Y/Z plot script covered later in this guide.
Step 9: Install ControlNet and Must-Have Extensions
ControlNet is the extension that turns Stable Diffusion WebUI from a prompt-only tool into something you can direct with pose skeletons, depth maps, and edge detection, and it remains the single most-installed extension across the entire AUTOMATIC1111-family ecosystem. Install it from the Extensions tab by searching for “sd-webui-controlnet” in the Available sub-tab, clicking Install, then restarting the WebUI from the Installed sub-tab.
After restarting, download at least one ControlNet model (OpenPose for figure posing, Canny for edge-guided generation, or Depth for scene composition) into the extension’s models folder. Three extensions worth adding alongside it: ADetailer for automatic face and hand correction, Ultimate SD Upscale for tiled high-resolution upscaling (the site’s 4K AI image upscaling guide compares seven different upscaling tools if you want a deeper look at that step specifically), and the built-in Civitai Helper for pulling checkpoints and LoRAs directly from a URL instead of manual downloads.
Step 10: Use img2img and Inpainting for Editing
The img2img tab takes an existing image as its starting point rather than pure noise, which is how most product photography and consistent-character workflows actually operate in practice. Upload a base image, set the denoising strength between 0.3 and 0.6 for edits that preserve the original composition, or closer to 0.7-0.85 for a heavier stylistic reinterpretation.
The Inpaint sub-tab lets you mask a specific region (a face, a background object, a logo) and regenerate only that area while leaving the rest of the image untouched. This is the local-generation equivalent of the editing features found in hosted tools, and it runs with zero per-image API cost once your checkpoint is downloaded.
Two settings inside the Inpaint sub-tab change results more than people expect. “Inpaint area: Only masked” restricts the model’s attention to the masked region plus a small padding border, which produces cleaner results for small fixes like correcting a hand or removing a stray object. “Inpaint area: Whole picture” gives the model more context about the surrounding scene, which helps for larger masked regions where lighting and perspective need to stay consistent with the rest of the image. There is no universally correct setting; small precise edits favor the masked-only option, while large structural changes favor whole-picture context.
Step 11: Add LoRA Models and Textual Inversion Embeddings
LoRA files are small adapter models, usually 50 to 300MB, that fine-tune a base checkpoint toward a specific style, character, or object without retraining the whole model. Drop downloaded .safetensors LoRA files into models/Lora, then reference them in your prompt with the syntax <lora:filename:0.8>, where 0.8 is the strength, typically kept between 0.6 and 1.0.
If you want to train your own LoRA from a small image set rather than downloading one, the site’s existing custom AI image LoRA training guide covers that process end to end, including the dataset preparation steps that most tutorials skip.
Textual inversion embeddings work alongside LoRAs but through a different mechanism: instead of adjusting the model’s weights, an embedding teaches the model a new token that represents a concept, style, or specific subject, triggered by including that token’s name directly in your prompt text. Embeddings are typically much smaller than LoRAs, often under 100KB, and go in the embeddings folder rather than models/Lora. Stacking a checkpoint, a style LoRA, and a subject embedding in the same prompt is a common combination for getting a specific character rendered in a specific art style without training a dedicated checkpoint for that exact combination.
Step 12-13: Enable the REST API and Set Up Maintenance
To automate generation from a script instead of clicking through the browser UI, add --api to your COMMANDLINE_ARGS line and relaunch. This exposes a local REST API at the same address the WebUI runs on, documented automatically at http://127.0.0.1:7860/docs, which lists every available endpoint including /sdapi/v1/txt2img and /sdapi/v1/img2img.
For ongoing maintenance, pull updates periodically with a plain git pull from inside the repository folder, and back up your models and extensions folders separately from the core code, since those are the files that take the longest to rebuild from scratch if something breaks during an update.
# Update the install
git pull
# Re-run the launcher so it reinstalls any changed dependencies
webui-user.bat
How Local Generation Costs Compare to Cloud APIs
The honest answer to “is local generation actually cheaper” depends entirely on volume and whether you already own a capable GPU. If you are buying hardware purely for this purpose, a GPU in the $400-900 range pays for itself against a cloud API only after generating tens of thousands of images, since most hosted services charge fractions of a cent to a few cents per image. If you already own a gaming GPU with 8GB or more VRAM, the marginal cost of local generation drops to essentially electricity, which makes the local route attractive almost immediately for anyone generating more than a handful of images a week.
| Approach | Upfront cost | Marginal cost per image | Best fit |
|---|---|---|---|
| Local Stable Diffusion WebUI | $0 if GPU already owned | Electricity only, roughly fractions of a cent | High-volume, repeated, or sensitive-content generation |
| Hosted API (GPT Image, Nano Banana, FLUX.2, etc.) | $0 | $0.005-$0.25 per image depending on model and quality tier | Low-to-medium volume, no GPU available |
| Dedicated GPU purchase for local generation | $400-$2,000+ | Electricity only after purchase | Sustained high-volume production use |
There is a middle path worth mentioning: nothing stops you from running Stable Diffusion WebUI locally for iteration and testing, then switching to a hosted API for final high-resolution renders that need capabilities your local checkpoint does not have, such as FLUX.2’s text rendering or Nano Banana Pro’s 4K output. The two approaches are not mutually exclusive, and the REST API covered in the next section makes it straightforward to swap between a local and hosted backend inside the same automation script.
Common Pitfalls When Installing Stable Diffusion WebUI
Most installation failures trace back to one of these mistakes, and catching them before you start saves an hour of debugging later.
- Installing the wrong Python version. The classic branch wants Python 3.11.9 specifically; Python 3.12 or 3.13 can break the dependency resolver during first launch.
- Cloning the wrong branch. Running a plain
git clonewithout--branch classicmay pull a different default branch with different requirements and a different Python target. - Forgetting “Add Python to PATH” on Windows. Without it, the launch scripts cannot find the Python interpreter and fail immediately with a cryptic error.
- Installing a checkpoint in the wrong folder. Checkpoints belong in
models/Stable-diffusion, not the repository root or a LoRA folder, and a misplaced file simply will not appear in the dropdown. - Running –xformers on an RTX 50-series card. xformers support lags behind new GPU architectures; Blackwell-generation cards need
--sageor default attention instead. - Skipping the GPU driver update. An old NVIDIA driver that only supports CUDA 11.x will silently fail against PyTorch builds compiled for CUDA 12.9.
- Filling the SSD before extracting. One checkpoint plus the base install can exceed 25GB quickly; running out of disk space mid-install corrupts the environment and usually requires a clean re-clone.
Troubleshooting: 8 Common Errors and How to Fix Them
These are the errors that show up most often in community support threads, along with the fix that resolves each one in the majority of cases.
- “CUDA out of memory” during generation. Add
--medvramor--lowvramto COMMANDLINE_ARGS, reduce batch size to 1, or lower the output resolution before upscaling separately. - “Torch is not able to use GPU.” Your driver is too old for the bundled PyTorch/CUDA build; update the NVIDIA driver, then delete the
venvfolder and relaunch to force a clean reinstall. - WebUI launches but the browser tab never loads. Check the terminal for the actual bound address and port; a firewall or another process may already occupy 7860, requiring
--port 7861as a workaround. - “No module named ‘xformers’” after adding the flag. xformers must be installed separately for some PyTorch/CUDA combinations; remove the flag temporarily or install the matching xformers wheel manually.
- ControlNet models not appearing in the dropdown. Confirm the model files sit in
extensions/sd-webui-controlnet/models, not the main checkpoint folder, and restart rather than reload the browser. - Extremely slow first generation, then normal speed after. This is expected; the first run compiles and caches attention kernels, and subsequent generations use the cached version.
- “RuntimeError: mat1 and mat2 shapes cannot be multiplied.” You loaded a LoRA trained for a different base architecture than your active checkpoint (for example, an SDXL LoRA on an SD 1.5 checkpoint); match the LoRA to the correct base model.
- Git pull fails with local changes conflict. Stash any manual edits to tracked files with
git stashbefore pulling, then reapply withgit stash poponce the update completes.
Advanced Tips for Speed, Memory, and Batch Automation
Once the basic install is stable, a handful of advanced settings meaningfully change throughput. Enabling --uv-symlink cuts the on-disk footprint of your Python environment from roughly 7GB down to around 100MB by symlinking shared packages instead of duplicating them, which matters if you run multiple WebUI installs side by side for testing different forks.
For batch work, the X/Y/Z plot script under the Scripts dropdown lets you sweep CFG scale, sampler, and step count in a single grid generation, which is far faster than manually testing combinations one at a time. If you are running on a headless server rather than a desktop, launch with --listen --api --nowebui to expose only the API without the browser interface, reducing memory overhead on machines dedicated purely to automated generation.
SageAttention, enabled with the --sage flag, typically cuts sampling time by a noticeable margin on RTX 40-series and newer cards compared to default attention, though actual gains depend on resolution and checkpoint architecture. Test it against xformers on your specific hardware rather than assuming one is universally faster.
If you plan to run generation across more than one GPU, Forge Classic does not have built-in multi-GPU load balancing the way a dedicated inference server would. The practical workaround most production setups use is running a separate WebUI instance per GPU, each bound to a different port with --port and pinned to its card with the CUDA_VISIBLE_DEVICES environment variable, then distributing requests across them from your automation layer with a simple round-robin queue. It is less elegant than a purpose-built inference cluster, but it works reliably and does not require touching the WebUI’s internals.
Model switching is another hidden cost worth planning around. Swapping checkpoints mid-session via the API or UI dropdown reloads several gigabytes from disk into VRAM, which can add 10-30 seconds depending on your storage speed. If a workflow needs to alternate between two checkpoints repeatedly, keeping both loaded is not possible in a single Forge Classic process, so batching all generations for one checkpoint before switching to the next avoids paying that reload cost over and over.
Complete Working Project: Automated Batch-Generation Script
With --api enabled from Step 12, the following Python script connects to your local WebUI instance and generates a batch of images from a list of prompts, saving each one with its prompt recorded in the filename. This is the kind of automation that replaces manually clicking Generate dozens of times for product catalogs, game asset variations, or marketing content drafts.
import requests
import base64
import os
from datetime import datetime
WEBUI_URL = "http://127.0.0.1:7860"
OUTPUT_DIR = "batch_output"
prompts = [
"a weathered leather backpack on a wooden table, studio lighting",
"a minimalist ceramic mug, soft shadow, white background",
"a vintage film camera on a marble surface, warm tones",
]
os.makedirs(OUTPUT_DIR, exist_ok=True)
for i, prompt in enumerate(prompts):
payload = {
"prompt": prompt,
"negative_prompt": "blurry, low quality, watermark",
"steps": 28,
"cfg_scale": 7,
"width": 1024,
"height": 1024,
"sampler_name": "DPM++ 2M Karras",
"seed": -1,
}
response = requests.post(f"{WEBUI_URL}/sdapi/v1/txt2img", json=payload)
response.raise_for_status()
result = response.json()
for j, image_b64 in enumerate(result["images"]):
image_data = base64.b64decode(image_b64)
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
filename = f"{OUTPUT_DIR}/batch_{i}_{j}_{timestamp}.png"
with open(filename, "wb") as f:
f.write(image_data)
print(f"Saved: {filename}")
print(f"Batch complete: {len(prompts)} prompts processed")
Run this with python batch_generate.py while the WebUI is running in another terminal window. Each prompt in the list produces one saved PNG in the batch_output folder, and you can extend the script to read prompts from a CSV file for larger production runs, or wrap the request loop in a queue system if you need to throttle GPU load across multiple concurrent jobs.
A few practical extensions turn this from a demo script into something production-ready. Wrapping the requests.post call in a retry loop with exponential backoff handles the occasional timeout that happens when a very large batch or a high step count pushes a single request past the default timeout window. Adding a time.sleep() call between requests prevents the script from saturating the GPU queue if you are running other processes against the same WebUI instance simultaneously. For teams generating product imagery at scale, logging each request’s prompt, seed, and resulting filename to a CSV or database alongside the saved PNG makes it possible to regenerate or audit any specific output later, which matters once you are producing hundreds of images a day rather than a handful for testing.
Frequently Asked Questions
Is Stable Diffusion WebUI free to use?
Yes. The software itself is open source and free; the only cost is the electricity and hardware to run it, unlike hosted APIs that charge per image.
Do I need an NVIDIA GPU, or will AMD or Apple Silicon work?
NVIDIA GPUs have the smoothest path because PyTorch’s CUDA builds are the most mature. AMD cards can run through ROCm on Linux with extra setup steps, and Apple Silicon Macs can run via the MPS backend, though both options generate images more slowly than an equivalent NVIDIA card.
What is the difference between Forge, Forge Classic, and reForge?
The original Forge by lllyasviel has not had a meaningful update since July 2025. Forge Classic and reForge are both actively maintained continuations; Forge Classic’s classic branch focuses on stability for SD 1.5 and SDXL, while reForge pulls in newer upstream AUTOMATIC1111 changes more aggressively.
Can Stable Diffusion WebUI run FLUX or Qwen-Image models?
Support is inconsistent across the AUTOMATIC1111-family forks. For FLUX, Qwen-Image, and other newer transformer-based architectures, ComfyUI has the most reliable and complete node support.
How much VRAM do I actually need to get started?
4GB is enough for basic SD 1.5 generation with the --lowvram flag. For comfortable SDXL use, 8-12GB is the practical minimum most users report.
Why is my first image generation so slow compared to later ones?
The first generation after launch compiles and caches attention kernels for your specific GPU and settings. Every generation after that reuses the cache and runs significantly faster.
Can I use checkpoints downloaded from Civitai?
Yes, as long as the file is in .safetensors or .ckpt format and placed in the models/Stable-diffusion folder. The .safetensors format is strongly preferred since it cannot execute arbitrary code on load, unlike older pickle-based .ckpt files.
Is it safe to run the REST API exposed to my local network?
The --api flag alone only binds to localhost by default. Adding --listen to expose it to other devices on your network should only be done behind a firewall or VPN, since the API has no built-in authentication.
Should I pick AUTOMATIC1111, Forge Classic, or reForge for a brand-new install in 2026?
For a first install, Forge Classic’s classic branch is the most balanced choice: it inherits AUTOMATIC1111’s extension ecosystem, keeps Forge’s memory-management improvements, and is actively maintained. Plain AUTOMATIC1111 is still a reasonable choice if a specific extension only supports it, and reForge is worth trying if you specifically want the newest samplers before they land elsewhere.
Can I run Stable Diffusion WebUI entirely offline after installation?
Yes, once Python dependencies, PyTorch, and your checkpoints are downloaded, the WebUI generates images with no internet connection required. The only time it needs network access afterward is to check for extension updates or pull new models.