Skip to content

Regarding the issue of training details. #4

Description

@Feynman1999

Thank you for your work.
You mentioned in the appendix that you directly shifted from a bidirectional 4-step model to self-forcing DMD training. I would like to ask:

  1. if this approach cause a temporal discontinuity in the visual effect (especially in the early iterations)? Since the model does not possess autoregressive (AR) capabilities and only has DMD loss, can we still achieve good results by shifting the model from bidirectional attention to unidirectional attention?
  2. Is batch size important? If we train with a batch size of 8, will the performance drop significantly?

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions