In some application scenarios, it is necessary to generate videos with two subjects. The current approach is to first use image editing to place both subjects in one frame, and then generate the video using the first frame as a control image. However, the ability to maintain ID consistency is weak and the results are not ideal. Will you consider supporting multi-image reference for better ID consistency?
In some application scenarios, it is necessary to generate videos with two subjects. The current approach is to first use image editing to place both subjects in one frame, and then generate the video using the first frame as a control image. However, the ability to maintain ID consistency is weak and the results are not ideal. Will you consider supporting multi-image reference for better ID consistency?