← LocalClaw BlogLocalClaw × OpenClaw × on-device decisions

One prompt.
The right model.

LocalClaw's new Routed Chat beta uses a small decision model on your Mac to choose between the chat models you mapped for simple questions, deeper analysis and code. The route is local, visible and reviewable.

OpenClaw requirementOpenClaw 2026.9.6 or newer
LocalClaw featureRouted Chat · beta, prepared for 1.0.209
Public download1.0.208 remains current until the signed release ships
LocalClaw Routed Chat showing local Simple, Code and Analysis decisions mapped to GPT-6 Luna, Sol and Astra
Product preview from the LocalClaw 1.0.209 beta branch. The screenshot shows three real routing paths; it is not a benchmark or a savings estimate.
01 / THE SHORT ANSWER

Stop asking one model to be everything.

A greeting, a strategic comparison and a Swift debugging request do not need the same kind of model. Yet most chat interfaces make you choose one model before you know what the next message will require—or leave every turn on the strongest, most expensive option.

Routed Chat turns model choice into a small, explicit decision before each reply. You map one available chat model to Simple, another to Analysis and one to Code. When you send a message, a local ONNX classifier proposes the route. LocalClaw applies a conservative policy, asks OpenClaw to use the mapped chat model and then verifies which model actually answered.

The router does not write the answer. It selects the chat model that will handle the complete request. Your ordinary Fast, Deep, Local and Cloud modes remain direct choices outside Routed Chat.
02 / THE FLOW

A decision on the Mac, then the full conversation.

Your messageThe current text plus a small amount of prior context for routing
Local ONNX decisionGLiNER 2.5 Small estimates Simple, Analysis and Code
Selected chat modelOpenClaw receives the complete prompt with the requested model

The setup downloads verified model artifacts once and prepares a dedicated localclaw-router agent. It does not replace the default decision model you may already use elsewhere in OpenClaw. After setup, the routing evaluation runs through OpenClaw's local ONNX provider on the Mac.

OpenClaw 2026.9.6 introduced the compatible Decision Model role and published the optional ONNX package. Its official plugin uses a persistent Node process and ONNX Runtime's CPU backend. OpenClaw documents that inference sends neither the supplied state nor the question to a remote decision service. Model artifacts are downloaded explicitly and checked against pinned revisions, sizes and hashes.

Sources: OpenClaw 2026.9.6 release notes · official ONNX decision-model guide.

03 / THE MODEL MAP

Three routes you control.

Simple

Routine work

Greetings, brief clarification and short translation can use the economical model you choose. In the recommended GPT-6 map, that slot is Luna.

Analysis

More reasoning

Planning, comparison, calculations and ambiguous requests use the stronger analysis model. The recommended GPT-6 map uses Astra.

Code

Software work

Programming, debugging and implementation requests go to the code model. The recommended GPT-6 map uses Sol.

Those GPT-6 choices are a preset, not a lock-in. Routed Chat reads the models available to the selected OpenClaw chat agent and lets you build your own map. The Code slot may reuse the Analysis model, but Simple and Analysis must differ so the router can make a meaningful choice.

The local model's score is not treated as unquestionable. LocalClaw checks for software terms, explicit analysis work, short translations and brief conversational messages. When the top category scores are close—or the request lacks enough context—the policy falls back to Analysis. That is deliberately conservative: uncertainty should not silently downgrade a task.

The decision model proposes. LocalClaw applies the policy. OpenClaw runs the selected chat model.
04 / VISIBLE BY DESIGN

No invisible router behind the chat box.

Every turn keeps the routing evidence next to the reply. Routed Chat shows the ONNX proposal, the final local route, the selected model, the model OpenClaw reports actually using and the relative category scores. If LocalClaw changed the proposal because of a conservative rule, the explanation says so.

What you seeWhy it mattersWhat it does not prove
Local routeSimple, Analysis, Code or an uncertainty fallbackThat the final answer will be correct
Relative category scoresHow the classifier ranked the three supplied labelsCalibrated answer confidence
Selected modelThe model LocalClaw asked OpenClaw to useThat the provider honored the request
Actually usedThe model identity returned by OpenClawProvider pricing or remaining quota

If the router is unavailable, the configured models are invalid or OpenClaw reports a different model than requested, LocalClaw surfaces the problem. It does not quietly pretend the route succeeded.

05 / WHAT STAYS LOCAL

Local routing is not the same as a fully local chat.

The decision evaluation runs on the Mac and does not consume cloud inference tokens. The model sees a bounded excerpt prepared for classification rather than taking over the conversation. Long prompts use their beginning and end for that local decision.

The complete message still goes to the selected chat model. If you map a cloud or OAuth model, its provider processes that chat request under the provider's account, privacy and billing rules. If you map a genuinely local chat model, that later inference can remain local too. Routed Chat does not make a cloud model private by placing a local classifier in front of it.

Local

The routing judgment

The ONNX classifier and LocalClaw's policy choose the task category on the Mac.

Depends on your map

The generated answer

The selected local, cloud or OAuth chat model receives the full request and produces the response.

06 / COST AND SPEED

A routing strategy, not a savings guarantee.

The economic idea is straightforward: do not send every “hello,” rewrite or short translation to the strongest model in the account. Reserve more capable models for the requests that appear to need them. But the result depends on your provider, subscription, token accounting, model map and actual workload.

Routed Chat therefore makes no universal percentage claim. A useful test compares total cost, latency and task quality with and without routing. Count wrong routes and manual retries, not only the price of the successful calls. The local classifier also has a small latency and memory cost of its own.

For OpenAI models, chat usage continues to follow the account or API billing path already configured in OpenClaw. The local decision itself does not consume OpenAI inference tokens.

07 / BETA BOUNDARIES

What this first version does—and does not do.

Included

Text routing with proof

Local classification, configurable model mapping, preview, conservative fallbacks and requested-versus-actual model reporting.

Not connected yet

Browser and general agent routing

The beta is a dedicated text conversation surface. It does not automatically route every OpenClaw channel, browser action or background agent job.

The current classifier accepts a limited input window. LocalClaw bounds the routing excerpt and keeps the full text for the selected chat model. Scores compare the supplied categories; they do not measure whether the answer is true, safe or complete.

Routed Chat also requires a local OpenClaw Gateway and OpenClaw 2026.9.6 or newer. Setup is explicit because it installs the optional ONNX provider, downloads verified model files and adds a dedicated router agent.

08 / THE BIGGER DECISION LAYER

OpenClaw supplies the primitive. LocalClaw turns it into a product flow.

OpenClaw's Decision Model role is provider-neutral. It can return a choice, score or Boolean probability through hosted Jev, a local Kev server or local ONNX classifiers. Simply selecting a decision model does not make OpenClaw route every chat prompt automatically.

Routed Chat is LocalClaw's explicit consumer of that primitive. It defines a three-way rubric, applies a product-level fallback policy, requests the mapped chat model and exposes the result in the interface. That distinction matters: the OpenClaw capability is general, while Routed Chat is one concrete workflow built on top of it.

For the architecture, providers and broader use cases, read our guide to OpenClaw, Jev and native Decision Models. For local research alternatives, see the visual comparison of Jev, Nimble and Laya.

09 / QUICK ANSWERS

LocalClaw Routed Chat FAQ

Does Routed Chat run the whole conversation locally?

No. The routing decision runs locally. The complete prompt then goes to the local, cloud or OAuth chat model selected by your map.

Can I choose my own models?

Yes. Assign available chat models to Simple, Analysis and Code. LocalClaw also offers a recommended GPT-6 map with Luna, Astra and Sol when those choices are available.

Will it always choose the cheapest model?

No. It chooses a task category, not a price. You decide which model represents the economical route. Ambiguous requests fall back to Analysis.

Are the category scores answer confidence?

No. They are relative estimates for the supplied routing labels. They do not predict whether the selected chat model's answer will be correct.

Is Routed Chat in the current download?

Not yet. The feature is prepared for LocalClaw 1.0.209 beta. The public installer remains 1.0.208 until the new app and DMG complete release verification.

LOCALCLAW'S TAKE

Model choice should become part of the interface.

A small local model can make a useful first decision.

Routed Chat does not try to replace the model that writes, reasons or codes. It gives that work a visible front door: classify the request locally, apply a cautious policy and ask the chat model best suited to the route.

The beta is intentionally narrow. Three categories are understandable. Every route can be inspected. Uncertainty has a stronger fallback. The next step is to measure real conversations and improve the rubric without hiding the trade-offs.

Explore the LocalClaw Mac app →

Primary sources and methodology

Sources and LocalClaw beta implementation checked September 24, 2026. The three screenshot examples demonstrate routing behavior; they are not a quality, latency or cost benchmark.