zgba 站群
Kev: Tiny Jev-like family of decision models built on top of Qwen3.5

Kev: Tiny Jev-like family of decision models built on top of Qwen3.5

Small Jev-like decision models you can train and run yourself.

Kev is a family of small decision models built on Qwen3.5 and based on the architecture described in Jev’s Architecture Unmasked. You can use the pretrained weights or train your own. The API matches TypeSafe’s System One, so you can point their Python SDK at your local server.

You’ll need Python 3.12+ and uv.

This starts Kev-4B locally. The first run downloads the adapter and base model. —run also accepts a local checkpoint directory or a Hub revision, such as jaredpalmer/kev-4b@qwen3 for the previous generation.

In another terminal, send it a ticket:

Example response from Kev-4B, running in bf16 on an Apple M5:

The ticket mentions a return, a late delivery, and a billing problem, and the department probabilities say so. That is the point of getting probabilities back instead of a single label.

The TypeSafe SDK is included in uv sync —extra serve:

With the server still running, open another terminal. You’ll need Node 20.9+:

Open localhost:3001, load a preset, and edit the text and questions. Press ⌘↵ to run it. “Packed vs separate” compares asking all questions at once with asking them one at a time. “Permute” runs a Choice question with six option orders. There are also presets for testing question isolation and fake delimiter tokens.

There’s a chess demo, too. The board is the input, legal moves are Choice options, and a Score question rates the position. You can play against Kev or let it play itself. Games are saved in localStorage.

Start with Kev-4B. Use Kev-9B when accuracy and calibration matter more than memory. Use Kev-0.8B if you need the smallest model. All three are built on Qwen3.5 bases with the same training data and settings.

Each cell is development / test. “Trained sources” means held-out examples from the datasets used to train Kev. “New sources” means datasets and policy rule types Kev wasn’t trained on. Every model was evaluated on the same development sets (decision-v7, transfer-v4) and the same test sets, which were read once per released checkpoint, after model selection. Lower Brier is better.

Kev-9B trails Jev by 3.5 points on the new-source development set (0.822 vs 0.857) and scores 0.852 on the test set, which Jev hasn’t been run on. We don’t know which datasets Jev was trained on, so this isn’t a controlled comparison of the two architectures.

All three models were updated on 2026-09-21 with a short second training pass on generated examples: policy cases with explicit day counts, and cases whose deciding evidence was removed, trained toward a uniform answer. On the test set this moved Kev-9B from 0.837 to 0.852 (95% CI +0.8 to +2.9 points), Kev-4B from 0.832 to 0.837, and Kev-0.8B from 0.668 to 0.684. The previous weights are at revision v7-base. Details and costs are in the model cards and PLAN.md.

Probabilities are calibrated by default. Each checkpoint stores a temperature (about 2.1–2.4) fitted on its in-distribution development set, and the pointer head applies it when the model is loaded. It never changes an answer: on new sources Kev-9B’s calibration error goes from 0.106 to 0.042 and its confident errors (wrong answers with probability ≥ 0.9) from 8.7% to 4.0%, about Jev’s 3.7%, with accuracy identical. Set KEV_TEMPERATURE=1.0 for the raw logits. The accuracy numbers in the table are the same either way; the Brier numbers are for the raw logits.

One optional setting: KEV_DATE_FACTS=1 appends the number of days between any two absolute dates found in the state (“June 26, 2026 is 8 days before July 4, 2026”). Kev can’t subtract dates reliably but it can use a stated day count: on the deadline policy questions Kev-9B goes from 0.80 to 0.90 (Jev 0.93). The table above doesn’t use it.

All weights are in the Kev collection and the GitHub release, which includes tarballs and SHA-256 checksums.

The first Kev family used Qwen3 bases with the same data and settings. Those weights stay published and are the faster choice on a Mac (see Serving Performance), but they are no longer developed.

Because only the base changed, the two generations are a controlled comparison. On the development set the accuracy gain is within noise; on the test set Kev-9B is 7.3 points ahead of Kev-8B (95% CI +2.8 to +11.7) with a Brier score 0.08 lower, Kev-4B is 2.9 points ahead of its predecessor (−0.9 to +6.4), and Kev-0.8B is 4.8 points ahead of Kev-0.6B (+0.2 to +9.3). PLAN_Qwen35.md has the full experiment, including the criteria we set in advance and how the results measured against them.

The original Kev-0.5B used Qwen2.5-0.5B and is kept for reference; see its model card.

state is the text to evaluate. Each question has instructions and, where needed, a set of answers to choose from.

For Choice with K > 1 options, confidence is (p_max − 1/K) / (1 − 1/K). A single option has confidence 1. Score confidence measures how close the distribution is to its most likely level. It’s an approximation of TypeSafe’s formula, which isn’t public. Neither field is a measured accuracy rate.

Objects and arrays are converted to labeled text. Delimiter-like strings in user input are escaped before tokenization. Invalid requests return 422. usage.output_tokens counts tokens in the serialized answers, not generated tokens.

The server binds to 127.0.0.1 and has no authentication. Keep it local unless you add authentication yourself.

Each checkpoint is a rank-16 LoRA adapter and a small pointer head on a Qwen base model. On an attention-only base (Qwen3), the state and questions go into one token sequence:

The attention mask lets a token read the state and its own question, but not other questions or future tokens. Each question’s position IDs restart just after the state. This lets the model process the state once and answer each question independently.

Qwen3.5 mixes attention layers with Gated DeltaNet layers, which are recurrent and ignore attention masks. For those models, each question runs as its own row: the state followed by that question, with the same positions as above. The rows are independent, so isolation is exact, and the server computes the state once and reuses its cache for every row. On attention-only models the two forms give identical probabilities (tests/test_model.py).

The pointer head scores each option’s hidden state against the question’s hidden state. A softmax turns those scores into probabilities. Because comes last, it can attend to the full option list.

Training uses cross-entropy on the correct answer. The adapter and head are trained together; the rest of the base weights stay fixed. Training examples and API requests use the same text format. No Jev outputs were used for training.

Asking questions together or separately produces probabilities within 4e-6 in the fp32 tests. This does not mean option order is irrelevant: options within a question can still affect one another. See the model code and parity tests.

On CUDA and ROCm, install flash-linear-attention for the Qwen3.5 models (the Modal image does this); a five-question request takes tens of milliseconds on an H100 and MI300X.

On Apple Silicon there are no fast kernels for the DeltaNet layers, so PyTorch runs reference code. Median model time in bf16 on an M5, five questions with three options each on a ~230-token state:

If you serve on a Mac and need low latency, use the Qwen3 models for now. An MLX backend for the Qwen3.5 models is the next planned change.

For the attention-only models the server merges the LoRA weights in fp32 before casting, uses SDPA attention on Apple GPUs, pads MPS inputs to 64-token buckets, and caches the state prefix for repeated requests (four states of at least 384 tokens by default). With a repeated 772-token state, Kev-4B (Qwen3) answers in 242 ms instead of 861 ms.

You can disable these with KEV_MERGE=0, KEV_ATTN=eager, KEV_SHAPE_BUCKET=1, and KEV_PREFIX_CACHE=0. On 24 new-source records, bf16 probabilities differed from fp32 by at most 0.017, with no change in the highest-probability answer. That is a small check, not a guarantee for every input.

The released models use decision-v7: 10,000 examples from ten public datasets, 896 generated policy examples, and 1,680 examples from 60 generated rule structures. All train for two epochs with LoRA rank 16 and cross-entropy. The learning rate is 1e-4 for 0.8B and 5e-5 for 4B/9B. For Qwen3.5 bases the adapter also covers the DeltaNet projections; kev.train picks the right targets from the model config.

The released models were trained on public datasets and generated policy examples. If your questions look different — your own routing categories, your own escalation rules, another language — a short fine-tune on a few hundred labelled examples usually helps more than any prompt change.

Put your examples in a JSONL file, one request per line. It’s the same shape as an API request, plus a label on every question:

For choice the label is the option name, for noul it’s true or false, and for score it’s the level’s position starting at 0. Keep 10–20% of the file aside for evaluation.

Then start from a released checkpoint with —init_from:

—init_from loads the adapter and pointer head from the released model before training, so you keep what Kev already knows and add your domain on top. Starting from the base model instead throws that away: in one user’s test on 836 support-tool decisions, a fine-tune from the base scored 0.33 on Kev’s own evaluation set, against 0.84 for the released model; the same data with —init_from kept 0.83 there and reached 0.88 on the new domain. Use a smaller learning rate than the from-scratch recipe (2e-5 is a good start), and pick —base to match the checkpoint you start from; the trainer checks that the base, revision, LoRA rank, and head size agree before it loads anything.

—batch 1 —accum 8 in bf16 fits the 0.8B model on a 4 GB GPU. The benchmark reports accuracy, Brier score, and calibration per question type, so you can see which of your questions the fine-tune helped. The checkpoint you started from is recorded in runs/mine/training_config.json.

Use uv run python -m kev.train —help for all training options. The released models don’t use the optional —perm_kl or —ord_w losses. The model cards have the training settings and dataset lists; PLAN.md records what was tried and what helped.

On a Mac, run one training job at a time. Two jobs on the same Apple GPU are much slower. Use Modal for longer runs.

Each trial gets its own H100. The study keeps running if you disconnect, and you can download the results when it finishes:

Study plans list training settings. Each trial saves the settings, code hashes, dataset hashes, and results. Choose models using the development results, not the locked test. After choosing a final candidate, you can read its test results once:

The evaluation data under evals/ is frozen: dataset versions and file checksums are recorded in each manifest. Large training files are downloaded from the Hub mirror and checked against those hashes.

These commands use development data. Test data requires —allow-test. The benchmark reports accuracy, Brier score, calibration error, the share of decisions you could automate at a 5% error budget, option-order changes, and question isolation. transfer-v9 adds 10-way MMLU-Pro, records buried among unrelated text, and “unknowable” records whose deciding evidence was removed; for those it reports how often the model still answers with at least 0.9 confidence (Kev-9B 5%, Jev 9%, Kev-8B 26%). Published accuracy numbers use fp32 evaluation, not the bf16 serving path.

evals/external/ holds two other projects’ test sets converted to this format, with their published live Jev results: SemIf’s 144 authored decisions (Kev-9B 0.917, Jev 0.965) and scienthoon’s 900 support tickets (Kev-9B 0.952 on routing and 0.911 on tone, Jev 0.897 and 0.914).

kev.jev runs the same questions against Jev through Vercel AI Gateway. kev.compare compares two saved runs with paired bootstrap confidence intervals. For the full experiment history, see PLAN.md and the leaderboard.

The API tests run TypeSafe’s example requests and the official SDK against your local server.

Built with Devin. Thanks to Archer Hume for the architecture write-up, TypeSafe for the API design, Qwen for the base models, 3x3xX3N0N for showing where the date-arithmetic failure really is, and Radexito for —init_from.

Related work: Hydragen, DeFT, FIRST.

Apache-2.0. The Qwen3 and Qwen3.5 base models are also Apache-2.0. Training datasets have their own licenses; see the model cards.

tiny Jev-like family of decision models built on top of Qwen3.5 you can train and run on your own

View original article