← All AI tools

Local AI Chat

Chat locally

Pick a compact model, wait for the one-time download, then chat. Inference runs on your device with WebGPU when available.

Runs in your browser. Model weights may download once from Hugging Face and cache locally — your prompts and files are not uploaded to BrowserSpaces.

On-device

Models run in your browser — chats stay local.

WebGPU

Uses your GPU when the browser supports it.

Small models

SmolLM2 and Qwen options sized for the web.

System prompt

Steer tone and behavior before you chat.

Use the tool

What's on your mind? Load a model, then ask a question.

Experimental. AI can make mistakes. Always verify important information. First model load downloads from Hugging Face; later visits reuse the cache.

Choose a model

Inference runs locally. First load downloads weights from Hugging Face into this browser's cache.

Stable

Top tier · WebGPU

Pick a model and load it to start.

A private assistant in the tab

Weights download once from Hugging Face and cache in the browser. Generation uses Transformers.js (ONNX Runtime) on WebGPU or WASM — your messages never hit our servers.

Before you start

  • Chrome or Edge with WebGPU is fastest.
  • Larger models need more RAM — close other heavy tabs.
  • AI can be wrong. Double-check facts and personal details.

Three steps

  1. 1

    Choose a model

    Start with SmolLM2 360M if you are unsure.

  2. 2

    Load

    First load downloads weights; later visits reuse the cache.

  3. 3

    Chat

    Ask questions, draft text, or brainstorm — verify important facts.