Skip to guide
Practical guide · September 21, 2026

Run an LLM.
Right in your browser.

Small models, local computation, no API key.

The short answer

Yes, a website can run a language model on your own computer. The page downloads model weights and a local inference engine, then your browser does the computation. LocalClaw Labs puts that into a free playground: chat, writing, a detective game and a decision comparison, with five model presets.

Open the browser AI playground →

What happens when you click Load model?

  1. The page loads the wllama engine, which brings llama.cpp inference into the browser.
  2. Your chosen GGUF model downloads from its pinned Hugging Face revision. This is a real download, from 397 MB to 3.01 GB in Labs, not a remote chatbot disguised as a local one.
  3. The browser stores the weights when local storage is available and prepares them for inference. Reloading a cached model still takes initialization time.
  4. You submit a prompt. WebAssembly runs the engine, with WebGPU acceleration when available. Smaller presets can use the CPU; Labs requires WebGPU for the 4B preset.

No model downloads automatically on page load. A slow connection can make the first setup take minutes, while browser storage limits or memory pressure can prevent a model from loading.

Choose the download that fits your device

Five downloadable browser presets · Apache 2.0 weights · fixed revisions
Model / sourceQuantizationDownloadMemory guidance
Qwen3 0.6BQ4_K_M397 MBAllow about 1.5 GB of free memory.
Qwen3 1.7BQ4_K_M1.11 GBAllow about 3 GB of free memory. Better suited to a computer.
Qwen3 0.6B Q8Q8_0639 MBAllow about 2 GB of free memory. Exact SemIf small-model build.
MiniCPM5 2BQ4_K_M1.56 GBAllow about 4 GB of free memory. SemIf’s desktop default.
Qwen3.5 4BQ4_K_M3.01 GBHigh-memory computer: allow about 6 GB free memory. WebGPU required; not recommended for phones.

Start with Qwen3 0.6B Q4 when testing compatibility. Move up only after it runs comfortably. The Q8 version uses more storage than the Q4 version of the same 0.6B model; the number of parameters alone does not determine download size. Qwen3.5 4B is a computer-oriented preset, not a sensible default for a phone.

Memory guidance is approximate. Available RAM, graphics buffers, context length, browser version and competing tabs all matter. A successful run on one machine does not prove that a model works on every browser or device.

Try something concrete before judging the model

In the detective game, ask each suspect about their alibi and compare the dialogue with the authored case notes. In Remix studio, rewrite the same text in several styles. In free chat, try a short question in your own language and watch the locally generated response.

For Decisions / RLCD, use Shuffle situation to explore 54 fictional scenarios. Support tickets, email checks and Detective Crab puzzles include both clear cases and deliberately incomplete evidence. Choose Noul for a yes/no estimate, Score for an ordered rubric, or Choice for unordered categories. There are 18 examples for each type. Browse the list or shuffle within a topic; neither action manufactures results. Edit the situation, question or rubric before running it. Noul keeps Yes/No outcomes fixed.

RLCD is a training objective, not a probability display

TypeSafe describes RLCD as Reinforcement Learning for Calibrated Decisions. Its Jev system should not be confused with a standard language model asked to pick a letter. SemIf, formerly OpenJEV, is an independent browser experiment; Labs includes its three open-model presets, not Jev weights.

Labs compares two outputs from the same model. One constrained output token provides relative scores over the allowed options. A second call asks the model to write probability estimates as JSON. The numbers can disagree, and the JSON can be malformed. Neither method establishes calibration: a displayed 80% is not evidence that the answer is correct 80% of the time.

Timings are measured on your current device. Warm-up is excluded; option scoring runs before JSON; measured prompt caching is disabled. The output instructions and token counts differ. Treat this as an experiment, not a controlled model ranking, a Jev benchmark or a tool for consequential automated decisions.

What stays local, and what still uses the network?

Prompts and answers in Labs stay in the tab and are not sent to an AI service. The page does not use session recordings. Hosting providers still receive ordinary requests for the page, runtime and weights. Reloading clears the conversation, but model files can remain cached. The privacy policy explains the wider website context.

Unload model releases active model memory. Remove downloaded models, inside How local inference works, removes the listed cached weights. A browser may also evict them to reclaim space. Inference after loading does not require an AI API, but the app does not guarantee offline page reloads. Do not equate local computation with a complete offline-installation guarantee.

Browser playground or desktop runtime?

A browser playground is useful for a quick demo, a lesson, a prompt experiment or a small interactive game. If you need larger models, long conversations, file workflows, persistent tools or runtime tuning, explore the local AI software directory and the full LLM catalogue. The free Labs page does not require the paid LocalClaw Mac application.

Shuffle a situation and try it →