Changing the model? Reload the page and pick a different one before loading.
A chat assistant running 100% in your browser — no server, no API key, fully private. Powered by WebGPU.
The model is downloaded once (a few hundred MB to ~2 GB) from a free CDN, cached in your browser, then runs on your own GPU. The first load is the slow part; after that it is instant.
Starting…
This demo needs WebGPU to run the model on your graphics card. Try:
--enable-unsafe-webgpu.Changing the model? Reload the page and pick a different one before loading.