This repository contains only the code and frozen configurations used to train and evaluate a controllable local-motif-to-Wagara diffusion model. Raw Wagara images, pretrained models, learned checkpoints, generated benchmark images, and participant responses are not distributed in this repository.
The model learns
P: representative local motif patch
c: full-pattern content caption
q: explicit repeat-structure vector
|
v
Y: complete Wagara surface pattern
p(Y | P, c, q)
q contains repeat mode, density, position irregularity, rotation variation,
scale variation, motif variation, axial orientation, and spacing anisotropy.
The target image, crop coordinates, masks, anchor canvases, and repeat maps are
not denoiser conditions.
Included:
- Phase 1 SDXL Wagara-prior LoRA training;
- Phase 2 motif, caption, and repeat-structure conditioned training;
- the structure-only
Pretrained IP+qbaseline; - paired neutral-q motif fidelity evaluation;
- counterfactual caption evaluation;
- q10-q90 repeat-structure evaluation; and
- blinded human-evaluation material generation and analysis.
Excluded:
- dataset collection, Qwen captioning, and motif mining;
- raw or copyrighted Wagara images;
- pretrained SDXL, IP-Adapter, DINOv2, and BLIP-2 weights;
- learned checkpoints and full generated-image collections; and
- manuscript drafts and private participant data.
This code expects prepared JSONL manifests. Paths may be absolute or relative to the execution directory.
| Stage | Train | Validation | Test | Total |
|---|---|---|---|---|
| Phase 1 full Wagara sources | 1,647 | 213 | 178 | 2,038 |
| Phase 2 motif-full pairs | 1,629 | 212 | 175 | 2,016 |
Each eligible source contributes one representative motif patch. Split
assignment is source-group preserving. See manifests/example.jsonl for the
minimum Phase 2 record shape.
SDXL, the VAE, and both text encoders are frozen. A rank-16 LoRA is trained on full Wagara images and content captions.
| Setting | Value |
|---|---|
| Epochs | 60 |
| Optimizer steps | 1,020 |
| Sample presentations | 98,820 |
| H200 batch ceiling | 104 |
| Learning rate | 5e-5 |
| Min-SNR gamma | 5 |
| Checkpoints | 25%, 50%, 75%, final |
The Phase 1 prior and all pretrained weights remain frozen. Phase 2 trains a rank-8 LoRA on IP image-attention K/V projections and a zero-initialized Fourier-feature structure encoder added to the SDXL timestep embedding.
| Setting | Value |
|---|---|
| Epochs | 80 |
| Optimizer steps | 1,360 |
| Sample presentations | 130,320 |
| H200 batch ceiling | 104 |
| Motif LoRA learning rate | 5e-5 |
| Structure learning rate | 2e-4 |
| Checkpoints | 25%, 50%, 75%, final |
Patch, caption, and q dropout are each 10%; all-condition dropout is 5%. Frozen weights use BF16 and trainable weights use FP32.
Python 3.10 and CUDA 12.4 were used for the locked run.
bash scripts/create_env.sh
source .conda/bin/activatePlace or link the immutable inputs at the paths configured in
configs/base.yaml:
models/pretrained/sdxl-base-1.0/
models/pretrained/ip-adapter/
models/pretrained/dinov2-small/
The public default uses physical GPU 0. To select another device, set both
CUDA_VISIBLE_DEVICES and runtime.visible_gpu to the same physical ID.
CUDA_VISIBLE_DEVICES=0 .conda/bin/python -m motif2wagara.cli train \
--config configs/phase1.yaml \
--set runtime.visible_gpu=0
CUDA_VISIBLE_DEVICES=0 .conda/bin/python -m motif2wagara.cli train \
--config configs/phase2.yaml \
--set runtime.visible_gpu=0
CUDA_VISIBLE_DEVICES=0 .conda/bin/python -m motif2wagara.cli train \
--config configs/pretrained_ip_q_baseline.yaml \
--set runtime.visible_gpu=0The checked-in batch size reproduces the H200 run. For a smaller GPU, override
train.batch_size and train.gradient_accumulation_steps together and report
the resulting effective batch and optimizer-step count.
Evaluation measures adherence to the actual inputs (P, c, q), not
reconstruction of the source full Wagara.
- 175 test motifs and four paired seeds;
- a neutral q to avoid rewarding or penalising intentional deformation;
- multi-scale local windows and rotation-compensated DINOv2 retrieval; and
- R@1, R@5, MRR, and repeat coverage.
The motif, q, diffusion seed, sampler, and non-target text are held fixed while one caption field changes. Background colour and textile texture are evaluated with atomic BLIP-2 VQA; background movement is also checked in CIELAB space.
For 31 held-out sources, motif, caption, noise, and all non-target q fields are held fixed. One of four structural attributes is changed from the training distribution's 10th to 90th percentile. Human or independent image-based judgements test whether the output changes in the intended direction.
Detailed contracts are in docs/fair_evaluation_protocol_v2.md and
docs/final_human_evaluation_protocol.md.
| Method | R@1 | R@5 | MRR | Coverage |
|---|---|---|---|---|
| Full | 62.57% | 85.86% | 0.731 | 0.784 |
| No Phase1 | 49.43% | 78.14% | 0.627 | 0.700 |
| Pretrained IP+q | 53.57% | 80.43% | 0.658 | 0.695 |
| No structure | 51.00% | 79.57% | 0.634 | 0.688 |
| Phase1+pretrained IP | 54.00% | 81.00% | 0.656 | 0.684 |
| Method | Attribute | Contrastive success | Strict joint success |
|---|---|---|---|
| Full | Background | 66.86% | 44.57% |
| No Phase1 | Background | 64.57% | 37.86% |
| Full | Texture | 60.44% | 19.48% |
| No Phase1 | Texture | 50.29% | 18.60% |
CIELAB background-direction success was 65.40% for Full and 61.59% for No Phase1.
| Attribute | Accuracy |
|---|---|
| Repeat density | 61.29% |
| Position irregularity | 67.87% |
| Rotation variation | 65.65% |
| Spacing anisotropy | 65.81% |
Machine-readable versions are under results/.
- Pretrained weights are immutable; learned weights are always written below
outputs/. - All result-producing manifests record source ID, split, seed, effective conditions, checkpoint paths, and output hashes.
- Only methods receiving the same effective conditions should be compared in a primary condition-fidelity table.
- Human responses are pending and must not be fabricated or inferred from automatic metrics.
Raw data and checkpoints are not hosted in Git. Add dataset provenance, licensing, checkpoint download URLs, and SHA-256 values here before archival release.