I would like to know how the Wan model is utilized to train multi-view tasks during Stage 1 and Stage 2.
Specifically, does the implementation follow a approach similar to FrameCrafter? For instance, do you discard the temporal dimension, and only encode/decode individual frames or views independently?
Any insights or high-level explanations regarding the adaptation of the spatial-temporal attention to multi-view consistency would be greatly appreciated!
I would like to know how the Wan model is utilized to train multi-view tasks during Stage 1 and Stage 2.
Specifically, does the implementation follow a approach similar to FrameCrafter? For instance, do you discard the temporal dimension, and only encode/decode individual frames or views independently?
Any insights or high-level explanations regarding the adaptation of the spatial-temporal attention to multi-view consistency would be greatly appreciated!