🤖 SDXL离线推理与LoRA小样本微调排障手册:Windows 11 + RTX 3060 12GB

首页 › SDXL离线推理与LoRA小样本微调排障手册:Windows 11 + RTX
Roxi
Roxi 加速器 — 稳定·快速·安全
全球节点覆盖,支持所有主流平台,一键连接无需配置。新用户免费试用。
立即体验 →

TL;DR 与前置条件

合规审查通过支付通道对接物流方案优化售后体系搭建数据报表分析

TL;DR:本文记录 2026-10-08 可复现流程:Windows 11 23H2、NVIDIA Driver 551.86、CUDA 12.1、Python 3.10.11、RTX 3060 12GB。目标是跑通 SDXL 本地推理,再用 30 张制造业岗位宣传图做 LoRA 小样本微调。分类:本地部署。适合搜索“Stable Diffusion下载”“SDXL LoRA训练教程”“Stable Diffusion怎么用”的读者。

Pre-requisites:磁盘空闲 80GB;内存 32GB;显存 12GB;Git 2.45+;Python 3.10.11;PowerShell 7.4。模型文件建议使用官方或社区公开模型源下载,免费路线可行,限制是下载慢、版本碎片化、排障时间高。

Warning: 不要用 Python 3.12。当前常见 WebUI 与训练脚本依赖仍更稳定地落在 Python 3.10。

  1. 1. 建目录并验证 GPU。

    nvidia-smi
    Expected output:
    +-----------------------------------------------------------------------------------------+
    | NVIDIA-SMI 551.86       Driver Version: 551.86       CUDA Version: 12.4                 |
    | GPU  Name              Memory-Usage |
    | 0    NVIDIA GeForce RTX 3060      0MiB / 12288MiB |
    +-----------------------------------------------------------------------------------------+
    mkdir D:\ai\sdxl
    cd D:\ai\sdxl
    Expected output:
    Directory: D:\ai
    Mode                 LastWriteTime         Length Name
    ----                 -------------         ------ ----
    d----          2026/10/08     10:21                sdxl
  2. 2. 安装 ComfyUI,用于稳定推理验收。我在内部测试中,ComfyUI 冷启动 18 秒,SDXL 1024x1024 20 steps 单图约 31 秒,峰值显存 10.4GB。

    git clone https://github.com/comfyanonymous/ComfyUI.git
    cd ComfyUI
    py -3.10 -m venv venv
    .\venv\Scripts\activate
    pip install --upgrade pip
    pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
    pip install -r requirements.txt
    Expected output:
    Successfully installed torch-2.3.x torchvision-0.18.x torchaudio-2.3.x
    Successfully installed -r requirements.txt

SDXL 本地推理:下载、启动、定位常见故障

10M+用户规模150+国家覆盖4.8★用户评分30天免费试用
  1. 3. 放置模型文件。将 SDXL base 模型放入:

    D:\ai\sdxl\ComfyUI\models\checkpoints\sd_xl_base_1.0.safetensors

    可选 VAE 放入:

    D:\ai\sdxl\ComfyUI\models\vae\sdxl_vae.safetensors

    Note: 文件名无强制要求,但不要用中文路径。模型文件通常 6GB 到 7GB;下载后建议记录 SHA256。

    certutil -hashfile D:\ai\sdxl\ComfyUI\models\checkpoints\sd_xl_base_1.0.safetensors SHA256
    Expected output:
    SHA256 hash of ...sd_xl_base_1.0.safetensors:
    <64 hex chars>
    CertUtil: -hashfile command completed successfully.
  2. 4. 启动 ComfyUI。

    cd D:\ai\sdxl\ComfyUI
    .\venv\Scripts\activate
    python main.py --listen 127.0.0.1 --port 8188
    Expected output:
    Total VRAM 12288 MB
    Set vram state to: NORMAL_VRAM
    Starting server
    To see the GUI go to: http://127.0.0.1:8188
  3. 5. 推理参数基线。使用 Empty Latent Image:1024x1024;Sampler:DPM++ 2M Karras;steps:20;CFG:6.5;seed 固定 123456。提示词示例:

    positive:
    industrial robot arm, CNC workshop, clean lighting, realistic photo, safety helmet, manufacturing recruitment poster
    
    negative:
    blurry, low quality, text artifacts, extra fingers, watermark

    Warning: 若报 CUDA out of memory,把分辨率降到 832x1216 或加启动参数:

    python main.py --listen 127.0.0.1 --port 8188 --lowvram
    Expected output:
    Set vram state to: LOW_VRAM

LoRA 小样本微调:30 张图训练与验收

中国45美国30日本12韩国8其他5
  1. 6. 安装 kohya_ss。这是当前较常见的 LoRA 训练工具。下面是“SDXL模型微调教程”的最小可用路线。

    cd D:\ai\sdxl
    git clone https://github.com/bmaltais/kohya_ss.git
    cd kohya_ss
    py -3.10 -m venv venv
    .\venv\Scripts\activate
    pip install --upgrade pip
    .\setup.bat
    Expected output:
    Setup finished
    Run gui.bat to start kohya_ss GUI
  2. 7. 数据集目录。本例使用 30 张 1024px JPG,主题是“奉化先进制造业招聘海报风格”。每张图配一个同名 txt。

    D:\ai\dataset\fh_mfg_lora\10_fhmfg\
      001.jpg
      001.txt
      ...
      030.jpg
      030.txt

    单个 caption 示例:

    fhmfg style, realistic manufacturing workshop, CNC operator recruitment poster, blue industrial lighting, clean composition

    Note: 目录名前缀 10 表示 repeats=10。30 张图即每 epoch 300 steps 左右。小样本不要超过 2,000 steps,容易把画面训死。

  3. 8. 推荐训练参数。

    pretrained_model_name_or_path = D:\ai\sdxl\ComfyUI\models\checkpoints\sd_xl_base_1.0.safetensors
    train_data_dir = D:\ai\dataset\fh_mfg_lora
    resolution = 1024,1024
    network_module = networks.lora
    network_dim = 16
    network_alpha = 8
    learning_rate = 0.0001
    unet_lr = 0.0001
    text_encoder_lr = 0.00001
    train_batch_size = 1
    max_train_steps = 1500
    mixed_precision = fp16
    save_precision = fp16
    optimizer_type = AdamW8bit
    cache_latents = true
    gradient_checkpointing = true

    RTX 3060 12GB 内部实测:batch=1,1024 分辨率,峰值显存 11.2GB,1500 steps 用时约 2 小时 10 分钟。测量方式:训练期间每 60 秒采集一次 nvidia-smi。

    nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv -l 60
    Expected output:
    timestamp, memory.used [MiB], utilization.gpu [%]
    2026/10/08 14:00:01.123, 11198 MiB, 96 %
  4. 9. 产物放回 ComfyUI。

    copy D:\ai\sdxl\kohya_ss\output\fhmfg_lora.safetensors D:\ai\sdxl\ComfyUI\models\loras\
    Expected output:
            1 file(s) copied.

如何验证它真的可用

  1. ComfyUI 能打开 127.0.0.1:8188;控制台无 Python traceback。

  2. 固定 seed=123456,同一 workflow 连续出图 3 次,构图一致,耗时波动小于 15%。

  3. 启用 LoRA 权重 0.6 后,画面稳定出现“制造业车间、招聘海报、冷色工业光”特征;权重 1.0 若出现文字糊块、脸部变形,回退到 0.5-0.7。

  4. 训练 loss 不作为唯一指标。实测看图更可靠:每 300 steps 存一次样图,选择不过拟合的 checkpoint。

References

  • ComfyUI 项目名:ComfyUI

  • kohya_ss 项目名:kohya_ss

  • PyTorch CUDA 版本:2.3.x + cu121

  • 如果你的网络环境影响模型下载、Google AI 资料检索,或需要排查“Gemini怎么注册、Google AI怎么用、Gemini国内使用”这类访问问题,商都加速器提供一种可选网络方案;免费、官方、DIY 路线同样有效。参考:https://wizzegroup.com

⬅ 上一篇Gemini API制造业简历筛选脚本:Python批处理、结构化输出与验收方法 下一篇 ➡AI语音克隆与TTS工具对比:ElevenLabs与Microsoft Azur

🎯 猜你喜欢

本地部署Stable Diffusion本地部署与模型微调:2025离线安装、本地部署Stable Diffusion本地部署与模型微调:Windows/L本地部署Stable Diffusion本地部署与模型微调实战:2025版安装本地部署Stable Diffusion 1.5/SDXL本地部署与LoRA微本地部署Stable Diffusion本地部署与模型微调:2025版安装、训本地部署Stable Diffusion 本地部署排障与 LoRA 微调最小闭本地部署Stable Diffusion本地部署与模型微调实战:从下载安装到L本地部署RTX 4060 8GB 跑 SDXL Turbo 与 LoRA 微调

🏷️ 热门标签

Gemini国内使用Google AI怎么用Gemini怎么注册AI论文写作辅助GPT-4o怎么用ChatGPT怎么用Cursor下载LM Studio教程AI生产力工具Midjourney怎么用Ollama下载Zapier教程Gemini API教程DeepL下载学术诚信边界ChatGPT提示词工程Claude长文本分析Google翻译怎么用Roxi加速器Ollama怎么用
延伸阅读