Skip to content

Audio Editing 实现细节 #37

Description

@Liyang-Chen-UCLA

感谢你们的工作,在Audio Generation中使用CoT很有启发性!

在论文4.4节 Stage3 Instruction-based Audio Editing 中,提到:
"The foundation model, conditioned on both this reasoning and the existing audio context, applies targeted modifications while maintaining overall coherence."

  1. 请问 audio context 指的是已经生成的audio,还是正处于生成过程中的audio latent ?

  2. 具体使用时,如何将 Audio 和 Edit Instruction 编码为condition,可以开源这部分的代码吗?

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions