September 20, 2026 · Research & analysis
Jev × Bespoke Nimble × LayaAI doesn’t always
need to write.
Sometimes, it needs
to decide.
Jev turns text into bounded judgments. You can’t currently download it. Nimble and the smaller Laya let you explore related ideas on your own machine.
Your app needs a route.
Not a paragraph.
A support message arrives. Before anyone writes a reply, the application needs to decide which team should handle it.
“I was charged twice for the same subscription.”
“This appears to be a billing issue because the customer mentions a duplicate charge…”
Your application still has to extract the decision.
The application receives an answer from a predefined set.
That is Jev’s proposition: give it context and explicit questions, then get structured judgments your software can use. TypeSafe calls this a “System One” model.
Sources: TypeSafe announcement · Typed questions
Which option?
Selects an option and returns probabilities and confidence.
How much?
Evaluates an ordered scale. The expected score can fall between levels.
How likely is “yes”?
Not a true/false value. No separate confidence field.
Jev is hosted.
Nimble can run locally.
Installing an SDK on your laptop does not put the model on your laptop. Follow where the input goes.
No public downloadable weights or documented self-hosted runtime identified as of September 20, 2026.
Download the model components first. Local inference can then keep your inputs on your machine.
Sources: Jev model documentation · Nimble repository · Published adapter
Score the allowed answers.
Skip the written response.
Nimble maps candidate answers to single-token codes, scores them, and normalizes their scores over the supplied choices. Ordinary code constructs the structured result.
B = technical
C = other
Read candidate-token logits
Code creates the object
Why this is interesting
A route, a flag, a category: many software decisions have a small answer space. Scoring that space can avoid generating an explanation nobody needs.
But the menu matters. If all your options are wrong, one still wins. An other option helps express uncertainty; it does not guarantee reliable abstention.
These add up to 100% among the supplied candidates. That alone does not mean the winning answer has an 80% real-world success rate.
Sources: Scoring design · Training approach
A valid answer can still
be the wrong answer.
Is this one of the allowed options?
Yes. The output follows the contract.
Does the message actually belong there?
That requires a separate evaluation.
Read Jev’s “zero hallucination” language through this distinction. A bounded output interface does not remove errors in understanding negation, indirect evidence or conflicting instructions. High confidence is not an independent fact-check.
Change one detail. Change the decision.
“The service is unavailable right now.”
The message explicitly says the problem is happening now.
Sources: Jev’s documented limitations · Confidence. Example design is LocalClaw’s.
Promising results.
Different strengths by task.
The Nimble project reports a newer evaluation covering 3,880 records across 13 public, human-labelled subsets. Here is agreement with the reference labels—not agreement between the two models.
Similar label agreement does not mean similar probability quality. Jev has lower reported calibration error in 11 of 13 subsets and lower Brier score in 10 of 13. Those measures examine how useful the probabilities are, beyond the winning label.
Source: Public benchmark report, pinned revision. Limitations include class imbalance, correlated examples and untested candidate-order sensitivity. Timings use different environments.
A smaller route to local decisions.
Laya is another open-source approach, with Apache-2.0 code and published weights. It uses a compact bidirectional encoder and a decision head, rather than Nimble’s 9B generative backbone. Its interface supports Choice, Score and Noul.
Choose a suitable checkpoint
322–421 million parameters
Returns typed answers
Sources: Official model card and architecture · Code licence.
ModernBERT-based. Default context: 512 tokens, including the question and options.
mmBERT-based. Default context: 1,024 tokens. The relevant starting point for French inputs.
A specialised checkpoint with 1,024-token context, fine-tuned for the evaluated workflows.
Can an ordinary computer run it?
It is a credible candidate for a consumer laptop or desktop. The loader explicitly handles CPU, NVIDIA CUDA and Apple MPS. CPU execution is supported, so a dedicated GPU is not required by the implementation. This is source-code verification, not a successful installation test on every platform.
Approximate FP32 parameter storage for 322–421M parameters. Loading copies, activations, Python and the operating system require additional memory.
A reasonable machine on which to try one checkpoint, based on size. No measured minimum is established here. Preloading all three uses more memory.
Implementation checked at agent.py, pinned revision. CPU/MPS use FP32 execution. RAM guidance above is an estimate, not a benchmark.
How would you start on a normal machine?
Use an isolated Python environment and the published pip install laya package. For a first French-language experiment, load only the multilingual checkpoint and explicitly choose CPU. These commands are illustrative and have not been run by LocalClaw:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install layaimport laya
agent = laya.load(
"convaiinnovations/laya",
subfolder="multilingual",
device="cpu",
)The shell example is for macOS/Linux; Windows uses a different environment activation command. The first load downloads model assets. The explicit subfolder limits checkpoint download selection in the inspected loader; a root load does not set that filter.
The advertised ~33 ms comes from a T4 GPU measurement, not a consumer-laptop test. The project’s own report says shipped probabilities can be overconfident and that its strong typed-decisions result comes from a specialised checkpoint. Its Jev comparisons use third-party results with different prompts or sample sizes. They do not prove that Laya is generally faster or better under identical conditions.
Source: Benchmark report. Laya is intentionally not added to the Jev/Nimble chart above: the evaluation protocols differ.
Let the model judge.
Let the application enforce.
Clear answer definitions
Candidate scores
Route or request review
Interpret language
Does the message describe an outage, a duplicate payment or an access problem?
Apply exact rules
Check dates, calculate amounts, enforce permissions and apply agreed routing priorities.
Resolve hard cases
Handle ambiguity and costly mistakes. Use observed errors to improve the workflow.
Before you automate
Define the policy
Specify labels, tie-breaks and the review path. A message can contain several issues.
Test your own examples
Compare against rules and a conventional model. Include negation and misleading instructions.
Measure the trade-off
Track errors as you automate more cases. Choose thresholds from evidence, not a demo score.
Local is possible.
It still needs real hardware.
Tokens, including context and scoring instructions, for the published checkpoint.
Qwen3.5-9B plus LoRA. The small adapter file is not the full model.
Weights alone, before runtime overhead. Not a tested minimum RAM requirement.
Sources: Repository and setup · Model card. Weight size is an estimate from parameter count and precision.
Apple Silicon and CUDA do not use the same scoring path
MLX’s ParallelScorer reuses a shared context prefix and batches field suffixes. The CUDA candidate scorer processes fields separately and repeats the full prefix. Do not assume identical performance characteristics. Quantized variants require their own compatibility and quality checks.
Read the implementation notes →What is different about the training?
TypeSafe describes Jev’s RLCD approach. Nimble uses supervised candidate classification with a LoRA adapter. Its synthetic, model-checked data includes contrastive examples: a small evidence change should flip the label. The original 2,676-example training set is distinct from the later 2,826-record collection.
Dataset documentation →Why not lead with Nimble’s 90.12% benchmark?
That result comes from a narrow, 324-example synthetic holdout; Jev scored 93.21% there. It is not a general production accuracy figure. The broader public suite above is more informative, but still cannot replace an evaluation on your workload.
Original evaluation →Does “local” mean every repository command stays offline?
No. Local inference can keep classification inputs on your machine after setup. Scripts that generate data using external models or compare results against Jev can make remote calls. Check the command you intend to run.
The idea worth taking home
Many applications need a handful of decisions before they need a sentence. Jev puts that idea at the centre of its interface. Nimble makes a related approach locally testable. Laya explores smaller encoders, making consumer-hardware experiments especially interesting.
Jev belongs on the LocalClaw reading list. Nimble and Laya belong on the local experimentation list.
Scope: public documentation checked September 20, 2026, including Jev 1.13.0. Nimble technical links are pinned where possible to revision 35fe1f4. No LocalClaw model benchmark was run. Diagrams are explanatory schematics, not internal architecture disclosures.