llm-web
Runs a tool-calling language model entirely in the browser, no server required.
An original Burn and wgpu implementation of the Qwen2 architecture, compiled to WebAssembly and running the full forward pass client-side with WebGPU: quantized GGUF weights, runtime LoRA adapters, schema-constrained decoding for tool calls, and a multi-step agent loop. Runs Qwen2.5-0.5B-Instruct in the browser with runtime LoRA adapters, the same engine behind llm-life’s language-model methods. The public demo runs SmolLM2-360M-Instruct.
- Runs
- browser, WASM, WebGPU
- Model
- Qwen/Qwen2.5-0.5B-Instruct, 0.5B params, ~430MB (Q4_0 GGUF), model licence Apache-2.0; code licence MIT
- Demo (GitHub Pages)
- https://idle-intelligence.github.io/llm-web/web/
- Source
- idle-intelligence/llm-web
- Status
- maintained
Learn more
Used on: llm, llm + tts, stt + llm + tts