Skip to content
#

sd35

Here are 7 public repositories matching this topic...

Language: All
Filter by language
WeeLLM

WeeLLM runs large diffusion models with as little as 4 GB of VRAM, without any quantization. It dynamically determines how many layers can fit within the available VRAM and streams the text encoder and transformer layers to the GPU layer by layer, enabling inference on hardware with limited VRAM. It supports both safetensors and GGUF models.

  • Updated Sep 10, 2026
  • Python

Add this topic to your repo

To associate your repository with the sd35 topic, visit your repo's landing page and select "manage topics."

Learn more