personaplex-24L-q4_k-webgpu
The pruned, QLoRA-recovered, Q4_K-quantized PersonaPlex-7B that sts-web's live demo actually loads.
nvidia/personaplex-7b-v1, by NVIDIA
Pruned, QLoRA-recovered, Q4_K quantization for WebGPU by Idle Intelligence.
NVIDIA Open Model License
NVIDIA PersonaPlex-7B-v1 with 8 middle temporal-transformer layers removed, quality recovered with a rank-32 LoRA trained on self-distilled teacher audio, then quantized to Q4_K for browser WebGPU inference. Quality assessed only by listening tests so far. Not affiliated with or endorsed by NVIDIA.
- Type
- model
- Runs
- WebGPU
- Size
- 6.74B params (24 of 32 temporal layers), 3.5GB (Q4_K)
- Demo (GitHub Pages)
- https://idle-intelligence.github.io/sts-web/web/
- Source
- idle-intelligence/sts-web
- HuggingFace
- idle-intelligence/personaplex-24L-q4_k-webgpu
- Status
- experiment
Learn more
Used on: sts