CLIP-G
Trained on a distillation set using HiDream CLIP-G as a base
Long prompts 100-200 word are compared in the vision/text pooled vector space and distilled to 77 token limit of SDXL
All training and distillation was at FP32
Trained on a distillation set using HiDream CLIP-G as a base
Long prompts 100-200 word are compared in the vision/text pooled vector space and distilled to 77 token limit of SDXL
All training and distillation was at FP32
0