Does anyone know if it is possible to use a smaller text encoder?
It is the bottleneck on my machine and I would assume that the actual music is generated by the diffusion model.
Maybe the text encoder could even be completely omitted if one knows how to prompt it correctly? At least I don't know how I would do that using ComfyUI.
I'd appreciate any feedback.
Does anyone know if it is possible to use a smaller text encoder?
It is the bottleneck on my machine and I would assume that the actual music is generated by the diffusion model.
Maybe the text encoder could even be completely omitted if one knows how to prompt it correctly? At least I don't know how I would do that using ComfyUI.
I'd appreciate any feedback.