Qingtao Pan, Kai Ye, Zhihao Dou, Bing Ji, and Shuo Li
ECCV, 2026
In this paper, we propose ELDiff, an evidential learning-supervised T2I diffusion model, which leverages the advantages of uncertainty metric and conflict detection to enhance the fault tolerance of unreliable segmentation maps and suppress semantic conflicts, strengthening object-wise consistency learning. Specifically, a pixel evidence loss is proposed to restrain overconfidence in unreliable labels through evidential regularization, and a token conflict loss is designed to weaken the contradiction between semantics through optimizing a measured conflict factor.
| Stable Diffusion Version | Checkpoint |
|---|---|
| v1.4 | ELDiff_SD14 |
| v2.1 | ELDiff_SD21 |
You can use the following code to download our checkpoints and generate images:
import torch
from diffusers import StableDiffusionPipeline
model_id = "./ELDiff_SD14"
device = "cuda"
pipe = StableDiffusionPipeline.from_pretrained(model_id, torch_dtype=torch.float32)
pipe = pipe.to(device)
prompt = "A cheese cat is sucking catnip"
image = pipe(prompt).images[0]
image.save("cat.png")Create and activate the environment:
conda create -n ELDiff python=3.8.5
conda activate ELDiff
conda install pytorch==1.13.1 torchvision==0.14.1 torchaudio==0.13.1 pytorch-cuda=11.7 -c pytorch -c nvidia
pip install -r requirements.txtcd train/data
# download COCO train2017
wget http://images.cocodataset.org/zips/train2017.zip
unzip train2017.zip
rm train2017.zip
bash coco_data_setup.shAfter this step, you should have the following structure under the train/data directory:
train/data/
coco_gsam_img/
train/
000000000142.jpg
000000000370.jpg
...
Download COCO segmentation data from coco_gsam_seg and put it under train/data directory.
After this step, you should have the following structure under the train/data directory:
train/data/
coco_gsam_img/
train/
000000000142.jpg
000000000370.jpg
...
coco_gsam_seg.tar
Then, run the following command to unzip the segmentation data:
cd train/data
tar -xvf coco_gsam_seg.tar
rm coco_gsam_seg.tarAfter the setup, you should have the following structure under the train/data directory:
train/data/
coco_gsam_img/
train/
000000000142.jpg
000000000370.jpg
...
coco_gsam_seg/
000000000142/
mask_000000000142_bananas.png
mask_000000000142_bread.png
...
000000000370/
mask_000000000370_bananas.png
mask_000000000370_bread.png
...
...
use the following command:
cd train
bash train.shThe results will be saved under train/results directory.
First, generate images for VISOR
cd evaluation/VISOR/generate_images
bash run_pipeline_VISOR.shSecond, calculate OA score
cd evaluation/VISOR/calculate_OA_score
python gen_img_obj.pycd evaluation/MULTIGEN
bash run_pipeline.shFirst, generate images for coco5k
cd evaluation/CLIP_score/generated_images/coco5k
bash run_pipeline_coco5k_multigpu.shSecond, calculate CLIP_score for coco5k
cd evaluation/CLIP_score/calculate_CLIP_Score
python text-image-similarity.pyFirst, generate images for flickr1k
cd evaluation/CLIP_score/generated_images/flickr1k
bash run_pipeline_flickr1k.shSecond, calculate CLIP_score for flickr1k
cd evaluation/CLIP_score/calculate_CLIP_Score
# modify the path: text_file_path = 'FlickrCaptions_1k.json' and generated_image = Image.open('/sample_imgs/flickr1k/CLIP-result/'+str(img_id)+'.png')
python text-image-similarity.pyFor FID calculation, we use the generated images from CLIP_score of coco5k and CLIP_score of flickr1k
pip install pytorch-fid
# calculate FID of coco5k
python -m pytorch_fid --batch-size=1 ./sample_imgs/coco5k/CLIP-result ./data/coco5k
# calculate FID of flickr1k
python -m pytorch_fid --batch-size=1 ./sample_imgs/flickr1k/CLIP-result ./data/flickr30kFor other evaluations such as T2I-CompBench, T2I-CompBench++, GenEval, and GenEval2, please follow their official code.
This repository is released under the Apache 2.0 license.
Our code is built upon diffusers, prompt-to-prompt, VISOR, Grounded-Segment-Anything, CLIP, and TokenCompose. We thank all these authors for their nicely open sourced code and their great contributions to the community.
If you find our work useful, please consider citing:
@misc{pan2026eldiffevidentiallearningmeets,
title={ELDiff: When Evidential Learning Meets Text-to-Image Diffusion},
author={Qingtao Pan and Kai Ye and Zhihao Dou and Bing Ji and Shuo Li},
journal={arXiv preprint arXiv:2606.20924},
year={2026}
}