llm-web

Runs a tool-calling language model entirely in the browser, no server required.

An original Burn and wgpu implementation of the Qwen2 architecture, compiled to WebAssembly and running the full forward pass client-side with WebGPU: quantized GGUF weights, runtime LoRA adapters, schema-constrained decoding for tool calls, and a multi-step agent loop. Runs Qwen2.5-0.5B-Instruct in the browser with runtime LoRA adapters, the same engine behind llm-life’s language-model methods. The public demo runs SmolLM2-360M-Instruct.

Runs
browser, WASM, WebGPU
Model
Qwen/Qwen2.5-0.5B-Instruct, 0.5B params, ~430MB (Q4_0 GGUF), model licence Apache-2.0; code licence MIT
Demo (GitHub Pages)
https://idle-intelligence.github.io/llm-web/web/
Source
idle-intelligence/llm-web
Status
maintained

Used on: llm, llm + tts, stt + llm + tts

← Home