@sierra youβre during vulkan inference and not cpu inference right?
@sierra it generates in real time for me and im on integrated graphics
@sierra cuda is probably the best option then. you could try the smaller Q4_K_M model quantization in the huggingface repo. how long is your reference audio? for decent results it doesnβt need to be longer than a minute
@sierra upload the file let me hear also i usually do 30s-1m so maybe more reference will make it better but also maybe slower
@sierra yeah i tried its okayish for me i think longer and more clear reference would make it better