personaplex-7b-v1-q4_k-webgpu
An unofficial quantization of a large speech-to-speech model for browser deployment.
nvidia/personaplex-7b-v1, by NVIDIA
Q4_K quantization for WebGPU by Idle Intelligence.
NVIDIA Open Model License
Unofficial Q4_K quantization of the full 32-layer NVIDIA PersonaPlex-7B-v1 (8.37B params, full-duplex speech-to-speech), shrunk from a 16.7GB bf16 checkpoint to a 4.4GB GGUF for browser WebGPU inference. Packaged for use with sts-web; not the layer-pruned variant the live demo actually loads (see personaplex-24L-q4_k-webgpu). Not affiliated with or endorsed by NVIDIA.
- Type
- model
- Runs
- WebGPU
- Size
- 8.37B params, 4.4GB (Q4_K)
- Source
- idle-intelligence/sts-web
- HuggingFace
- idle-intelligence/personaplex-7b-v1-q4_k-webgpu
- Status
- experiment