Autoregressive speech prediction with EnCodec and FACodec token language models: 2s context -> 1s future speech, with STOI/PESQ/DNSMOS analysis of where those metrics disagree.
pytorch speech-synthesis transformer language-model speech-processing pesq stoi encodec dnsmos neural-audio-codec facodec speech-quality-assessment
-
Updated
Aug 6, 2026 - Python