sts-web

A full-duplex speech-to-speech assistant running entirely client-side, still early and rough.

Browser-native speech-to-speech, 100% client-side via Rust/WASM + WebGPU. Runs a pruned 24-layer, Q4_K-quantized PersonaPlex-7B (QLoRA-recovered from NVIDIA’s 32-layer original) through a full-duplex pipeline: mic to Mimi encoder to temporal/depth transformers to Mimi decoder. Walkie-talkie mode works with a handful of voice presets; true full-duplex streaming is not supported yet, and audio quality is currently poor.

Runs
browser, WASM, WebGPU
Demo (GitHub Pages)
https://idle-intelligence.github.io/sts-web/web/
Source
idle-intelligence/sts-web
HuggingFace
idle-intelligence/personaplex-24L-q4_k-webgpu
Status
experiment

Used on: sts

← Home