Repository navigation
RuntimeError: The size of tensor a (4608) must match the size of tensor b (5120) at non-singleton dimension 2 during DreamBooth Training with Prior Preservation #10722
Description
Activity
And I guess this bug maybe is caused by following code, these code make the shape of text_ids from (512, 3) to (1024, 3). How can I fix it? Please help me
if not train_dataset.custom_instance_prompts: if not args.train_text_encoder: prompt_embeds = instance_prompt_hidden_states pooled_prompt_embeds = instance_pooled_prompt_embeds text_ids = instance_text_ids if args.with_prior_preservation: prompt_embeds = torch.cat([prompt_embeds, class_prompt_hidden_states], dim=0) pooled_prompt_embeds = torch.cat([pooled_prompt_embeds, class_pooled_prompt_embeds], dim=0) text_ids = torch.cat([text_ids, class_text_ids], dim=0)exactly same issue :(
And I guess this bug maybe is caused by following code, these code make the shape of text_ids from (512, 3) to (1024, 3). How can I fix it? Please help me
if not train_dataset.custom_instance_prompts: if not args.train_text_encoder: prompt_embeds = instance_prompt_hidden_states pooled_prompt_embeds = instance_pooled_prompt_embeds text_ids = instance_text_ids if args.with_prior_preservation: prompt_embeds = torch.cat([prompt_embeds, class_prompt_hidden_states], dim=0) pooled_prompt_embeds = torch.cat([pooled_prompt_embeds, class_pooled_prompt_embeds], dim=0) text_ids = torch.cat([text_ids, class_text_ids], dim=0)It seems that this code doesn't make the shape of '''prompt_embeds''', '''pooled_prompt_embeds''', and '''text_ids''' match the model_inputs. As the '''prompt''' is contained in every batch, I managed to solve this issue by simply modifying the line 1538-1546 to
else: elems_to_repeat = len(prompts) if args.train_text_encoder: prompt_embeds, pooled_prompt_embeds, text_ids = encode_prompt( text_encoders=[text_encoder_one, text_encoder_two], tokenizers=[None, None], text_input_ids_list=[ tokens_one.repeat(elems_to_repeat, 1), tokens_two.repeat(elems_to_repeat, 1), ], max_sequence_length=args.max_sequence_length, device=accelerator.device, prompt=args.instance_prompt, ) else: prompt_embeds, pooled_prompt_embeds, text_ids = compute_text_embeddings( prompts, text_encoders, tokenizers )This intended to encode the prompts in every training step to make sure that the shape of text embeddings match that of model_inputs
Reacted by yinguowei and QingshuiLYeah, I also find that because of the problem described in this issue, the text ids created no longer have the batch size dimension, so concatenating the class with the instance text ids in the following code results in a subsequent dimension error.
if not args.train_text_encoder: prompt_embeds = instance_prompt_hidden_states pooled_prompt_embeds = instance_pooled_prompt_embeds text_ids = instance_text_ids if args.with_prior_preservation: prompt_embeds = torch.cat([prompt_embeds, class_prompt_hidden_states], dim=0) pooled_prompt_embeds = torch.cat([pooled_prompt_embeds, class_pooled_prompt_embeds], dim=0) text_ids = torch.cat([text_ids, class_text_ids], dim=0)Since text ids are all 0 vectors (as set in the method), this bug can be fixed by not concat the text ids of class and instance, and the rest of the code is not a problem because python has a broadcast mechanism
Reacted by Chase Huh and QingshuiLReacted by Chase Huhgithub-actions commented
on Mar 13, 2025 on Mar 13, 2025 – with GitHub ActionsContributorMore actionsThis issue has been automatically marked as stale because it has not had recent activity. If you think this still needs to be addressed please comment on this thread.
Please note that issues that do not follow the contributing guidelines are likely to be ignored.
- addedstaleIssues that haven't received updatesIssues that haven't received updates
on Mar 13, 2025 text_ids = torch.cat([text_ids, class_text_ids], dim=0)causes incorrect concatenation. Removing this line will make it work properly.
Describe the bug
I am trying to run "train_dreambooth_lora_flux.py" on my dataset, but the error will happen if --with_prior_preservation is used.
Who can help me? Thanks!
Reproduction
python ./examples/dreambooth/train_dreambooth_lora_flux.py
--pretrained_model_name_or_path=$MODEL_NAME
--instance_data_dir=$INSTANCE_DIR
--output_dir=$OUTPUT_DIR
--with_prior_preservation
--class_data_dir="my_file"
--class_prompt="A photo"
--instance_prompt="A sks photo"
--resolution=1024
--rank=32
--max_train_steps=5000
--checkpointing_steps=100
--seed="0"
--mixed_precision="bf16"
--train_batch_size=1
--guidance_scale=1
--gradient_accumulation_steps=4
--optimizer="prodigy"
--learning_rate=1.
--report_to="tensorboard"
--lr_scheduler="constant"
--lr_warmup_steps=0
Logs
System Info
NVIDIA A800-SXM4-80GB, 81920 MiB
NVIDIA A800-SXM4-80GB, 81920 MiB
NVIDIA A800-SXM4-80GB, 81920 MiB
NVIDIA A800-SXM4-80GB, 81920 MiB
NVIDIA A800-SXM4-80GB, 81920 MiB
NVIDIA A800-SXM4-80GB, 81920 MiB
NVIDIA A800-SXM4-80GB, 81920 MiB
Who can help?
No response