Skip to content
jongwonryuPublic

About

Language-Grounded Multi-Domain Image Translation via Semantic Difference Guidance(EACL, 2026)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

LACE Training Code

This repository contains the training code for the proposed LACE framework: Language-grounded Attribute-Controllable Translation.

The core training entry point is dip_adapter.py. It trains the GLIP-Adapter with frozen Stable Diffusion, CLIP image encoder, and DINOv2 image encoder. The adapter concatenates a global CLIP image token with local DINOv2 tokens and injects them into the U-Net cross-attention layers.

Large datasets, pretrained weights, fine-tuned checkpoints, and generated figures are intentionally excluded from Git.

Key Files

  • dip_adapter.py: trains the proposed CLIP+DINOv2 image prompt adapter.
  • train.py: optional Stable Diffusion domain fine-tuning script.
  • ip_adapter/attention_processor.py: cross-attention processor with image tokens.
  • ip_adapter/resampler.py: resampler utilities used by adapter variants.
  • ip_adapter/utils.py: small generation/training utilities.
  • scripts/train_lace_adapter.sh: recommended adapter training command.
  • scripts/train_base_sd.sh: optional base Stable Diffusion fine-tuning command.

Data Format

Training data should be an image folder plus a JSONL metadata file:

{"file_name": "0.jpg", "text": "young woman, no glasses, big smile"}
{"file_name": "1.jpg", "text": "clear, city street, daytime"}

The metadata key can be either file_name or image_file.

Install

pip install -r requirements.txt
accelerate config

The original experiments used Stable Diffusion 2.1, CLIP-ViT-H/14, and DINOv2-Large. Put those model directories outside Git or under ignored paths.

Train The Adapter

Set paths with environment variables and run:

BASE_MODEL=/path/to/domain-finetuned-sd \
CLIP_ENCODER=/path/to/clip-image-encoder \
DINO_ENCODER=/path/to/dinov2 \
DATA_JSON=/path/to/metadata.jsonl \
DATA_ROOT=/path/to/images \
OUTPUT_DIR=/path/to/output/lace_adapter \
bash scripts/train_lace_adapter.sh

Important defaults in the script:

  • image resolution: 512
  • image prompt tokens: 257 (1 CLIP global token + 256 DINOv2 patch tokens)
  • mixed precision: fp16
  • learning rate: 1e-5

Optional Base Fine-Tuning

If a domain-specific Stable Diffusion checkpoint is needed before adapter training:

PRETRAINED_MODEL=stabilityai/stable-diffusion-2-1-base \
DATA_ROOT=/path/to/images-with-metadata-jsonl \
OUTPUT_DIR=/path/to/output/domain_sd \
bash scripts/train_base_sd.sh

For a minimal code release, committing dip_adapter.py, train.py, ip_adapter/*.py, scripts/, requirements.txt, .gitignore, and this README is enough.

About

Language-Grounded Multi-Domain Image Translation via Semantic Difference Guidance(EACL, 2026)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages