Local AI Chat
Chat locally
Pick a compact model, wait for the one-time download, then chat. Inference runs on your device with WebGPU when available.
On-device
Models run in your browser — chats stay local.
WebGPU
Uses your GPU when the browser supports it.
Small models
SmolLM2 and Qwen options sized for the web.
System prompt
Steer tone and behavior before you chat.
Use the tool
What's on your mind? Load a model, then ask a question.
Choose a model
Inference runs locally. First load downloads weights from Hugging Face into this browser's cache.
Stable
Top tier · WebGPU
How it works
A private assistant in the tab
Weights download once from Hugging Face and cache in the browser. Generation uses Transformers.js (ONNX Runtime) on WebGPU or WASM — your messages never hit our servers.
Before you start
- Chrome or Edge with WebGPU is fastest.
- Larger models need more RAM — close other heavy tabs.
- AI can be wrong. Double-check facts and personal details.
Three steps
- 1
Choose a model
Start with SmolLM2 360M if you are unsure.
- 2
Load
First load downloads weights; later visits reuse the cache.
- 3
Chat
Ask questions, draft text, or brainstorm — verify important facts.