personaplex-7b-v1-q4_k-webgpu

An unofficial quantization of a large speech-to-speech model for browser deployment.

nvidia/personaplex-7b-v1, by NVIDIA

Q4_K quantization for WebGPU by Idle Intelligence.

NVIDIA Open Model License

Unofficial Q4_K quantization of the full 32-layer NVIDIA PersonaPlex-7B-v1 (8.37B params, full-duplex speech-to-speech), shrunk from a 16.7GB bf16 checkpoint to a 4.4GB GGUF for browser WebGPU inference. Packaged for use with sts-web; not the layer-pruned variant the live demo actually loads (see personaplex-24L-q4_k-webgpu). Not affiliated with or endorsed by NVIDIA.

Type
model
Runs
WebGPU
Size
8.37B params, 4.4GB (Q4_K)
Source
idle-intelligence/sts-web
HuggingFace
idle-intelligence/personaplex-7b-v1-q4_k-webgpu
Status
experiment

← Home