adding more "hint" to training process

Author: orydatadudesCreated Mar 14, 2023Updated Nov 5, 2025

Hi, i was focusing with the human posture task (getting posture from openpose image + prompt and than generating the charter under the right pose - control_sd15_openpose.pth)

However, i wanted to add one more hint to force the controlnet to generate specific human: so if in the original code the hint be an posture image like that :

v2-c5e272899550ac318ed4732336fd7c82_720w

i would like to add more image of the specific human:

MEN_Denim_id_00000080_0_01_7_additional

the target should be that image of that person, under the new posture

so what i did is:

  1. in the dataset file: reading that extra image too, concatenate in the channel dimension, that image with the posture image so now the
    source variable is 6 channels not 3

     # concate source and source image
     source = np.concatenate([source,source_image],axis=2)
    
     return dict(jpg=target, txt=prompt, hint=source)
  2. changing the yaml config file to support 6 channels - NOT SURE I REALLY UNDERSTATED THE MEANING OF THESE VALUES

model: target: cldm.cldm.ControlLDM params: linear_start: 0.00085 linear_end: 0.0120 num_timesteps_cond: 1 log_every_t: 200 timesteps: 1000 first_stage_key: "jpg" cond_stage_key: "txt" control_key: "hint" image_size: 64 channels: was 4 i changed to 7 cond_stage_trainable: false conditioning_key: crossattn monitor: val/loss_simple_ema scale_factor: 0.18215 use_ema: False only_mid_control: False

control_stage_config:
  target: cldm.cldm.ControlNet
  params:
    image_size: 32 # unused
    **in_channels:  was 4 i changed to 7**
    **hint_channels: was 3 i changed to 6** 
    model_channels: 320
    attention_resolutions: [ 4, 2, 1 ]
    num_res_blocks: 2
    channel_mult: [ 1, 2, 4, 4 ]
    num_heads: 8
    use_spatial_transformer: True
    transformer_depth: 1
    context_dim: 768
    use_checkpoint: True
    legacy: False

unet_config:
  target: cldm.cldm.ControlledUnetModel
  params:
    image_size: 32 # unused
    **in_channels:  was 4 i changed to 7**
    **out_channels:  was 4 i changed to 7**
    model_channels: 320
    attention_resolutions: [ 4, 2, 1 ]
    num_res_blocks: 2
    channel_mult: [ 1, 2, 4, 4 ]
    num_heads: 8
    use_spatial_transformer: True
    transformer_depth: 1
    context_dim: 768
    use_checkpoint: True
    legacy: False

first_stage_config:
  target: ldm.models.autoencoder.AutoencoderKL
  params:
    **embed_dim:  was 4 i changed to 7**
    monitor: val/rec_loss
    ddconfig:
      double_z: true
      **z_channels:  was 4 i changed to 7**
      resolution: 256
      in_channels: 3
      out_ch: 3
      ch: 128
      ch_mult:
      - 1
      - 2
      - 4
      - 4
      num_res_blocks: 2
      attn_resolutions: []
      dropout: 0.0
    lossconfig:
      target: torch.nn.Identity

cond_stage_config:
  target: ldm.modules.encoders.modules.FrozenCLIPEmbedder

the problem is when i trained the model from scratch - running tutorial_train.py with resume_path = None the model predictions, the reconstruction and the samples that locate under image_log->train folder are just a noise

does anyone have any idea how to solve that ? thanks