Repo2 hours ago

Kev Gained 4,310 Stars in a Week and Was Never Wrong Above 0.6 Confidence on Our Archive

Jared Palmer's Kev gained 4,310 stars in seven days. Kev-4B sorted 97 of our 140 articles correctly, well short of Haiku's 123, but all 23 answers it gave at 0.6 confidence or higher were right.

The WJS Desk

Sep 30, 2026 · 7 min read

Photo by Timothy Huliselan on Pexels

Our trending tracker first saw jaredpalmer/kev at 16:00 UTC on 22 September, at 3,493 stars. By 13:07 UTC on 29 September it had 7,803, a gain of 4,310 in just under seven days. The fastest stretch was the first twelve hours we watched, when it added 1,056. It has cooled since then: the last 21 hours added 205.

That curve matches a Hacker News post on 21 September that reached 462 points and 210 comments, and the wider rush of Jev clones. Kev is one of several open "Jev-like" decision models in our tracker this month, alongside Laya and SemIf-OpenJev. We have already run one of the others, Laya, on our own archive. That gave us a test we could simply repeat.

What it actually is

Kev is a family of LoRA adapters on Qwen3.5 and Qwen3.8 base models, in four sizes: 0.8B, 4B, 9B and 27B. Like Jev, it answers typed questions about a piece of text in one pass (a choice from your labels, a score on a scale, or a noul yes/no probability) and generates no text. The server speaks TypeSafe's System One wire format, so the README says their Python SDK can point at it unchanged. We tested the HTTP API only, not the SDK.

uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b --port 8009

curl -s localhost:8009/v1/systemone -d '{
  "state": "Laya Added 24,000 Stars in Eight Days...",
  "questions": {"category": {"type": "choice",
    "instructions": "What kind of article is this?",
    "criteria": {"news": "...", "tutorial": "...", "repo": "...", "pulse": "..."}}}}'

The kev/ package is 8,862 lines of Python, but most of it is training: rounds.py alone is 1,363 lines. The Apple Silicon inference path in mlx_model.py is 154 lines and the server is 297. There are eight core dependencies, torch and transformers among them, and the serve extra adds FastAPI, uvicorn, the TypeSafe SDK and mlx-lm.

We ran it on our own archive

This is the test from our Laya piece: the same 140 articles, published before 28 September, each with a category a human assigned. We sent each title and excerpt to Kev-4B with the same four category definitions, one request per article, on an M4 Pro with 48 GB. The README recommends Kev-4B as the starting size, and it lists a 32 GB Mac as enough.

Setup was the slow part. The clone took 5 minutes 25 seconds, most of it a 161 MB git history. uv sync --extra serve installed 102 packages into a 1.1 GB virtualenv. The first server start took 811 seconds, 12 minutes of which went on downloading 9.3 GB of Qwen3.5-4B-Base weights and the 160 MB Kev adapter.

For the comparison we also reran Claude Haiku tonight, as one batched call through the Claude CLI with the same definitions. It scored 123 of 140. The Haiku run in our Laya piece scored 108 with a differently formatted prompt, which is a reminder of how much prompt shape moves an LLM baseline. The Laya row below comes from that earlier piece. We did not rerun Laya tonight.

140 of our articlesKev-4B (MLX)Laya 0.3.21Claude Haiku (tonight)
Correct category97 (69.3%)90 (64.3%)123 (87.9%)
News correct32 of 3917 of 3938 of 39
Tutorials correct27 of 5735 of 5747 of 57
Repo spotlights correct21 of 25n/a23 of 25
Time per article271 ms median45 ms136 s for one batched call
Download before first answer1.1 GB venv plus 9.5 GB weights714 MB venv plus 1.6 GB weightsNone

Kev beat Laya by seven articles, almost all of them news: 32 correct against 17. It did worse on tutorials. Twenty of our tutorials came back as pulse, mostly parts of our beginner prompting courses whose excerpts read like opinions. Our 271 ms median per article is below the README's 721 ms for five questions on an Apple M5, but we asked one question, so the two numbers are not comparable.

The README quickstart also ran, but did not reproduce exactly. On the sample ticket the README shows returns at 0.47 and escalation at 0.93. We got 0.51 and 0.72. The answers matched and the probabilities did not. That first request took 2,144 ms cold. Twenty repeats of the same text then took a 204 ms median, because Kev caches repeated states.

The calibration is the real feature

Kev reports a confidence with every choice, and on our data it was very conservative:

  • At 0.6 or higher: 23 answers, all 23 correct.
  • At 0.5 or higher: 36 answers, 35 correct.
  • Below 0.3: 59 answers, only 27 correct.

No answer went above 0.8825. By comparison, Laya was 49 of 52 correct at the 0.6 line. So Kev was right more often than Laya when it committed, but it committed much less often. We built the same cascade as last time, using Kev when it was at least 0.6 sure and Haiku for the rest. It scored 123 of 140, exactly what Haiku scored alone, and skipped Haiku on 23 articles. Against a strong Haiku baseline, Kev did not improve the answers. It removed 16% of the LLM calls without costing any accuracy.

The hidden cost: Kev-4B is not a small download. Before the first answer you need 9.3 GB of base weights, a 1.1 GB virtualenv pinned to torch<2.9 and Python 3.12 or 3.13, and a server process that holds a 4B model in memory. That is a lot of infrastructure for a model that, on our task, was confident enough to act alone on one article in six.

What it does not do

  • Long inputs. Training used states of up to 384 tokens. The README says accuracy falls on long documents, with Kev-9B scoring 0.556 on questions buried in 1k to 6k tokens of text.
  • Knowledge questions. The maintainers report MMLU of 0.74 for Kev-9B against 0.90 for Jev, and weaker date arithmetic on the smaller models.
  • Local fine-tuning, out of the box. The bundled fine-tuning skill runs the loop on Modal, which is a paid cloud GPU service.
  • Controlled comparison with Jev. The README says this itself: "We don't know what Jev was trained on, so this isn't a controlled comparison."
  • A container image. One HN commenter, webprofusion, asked why none of these projects ship one. Kev does not either, although it has a one-command Modal deploy.

Who made it, and does that matter

Jared Palmer, whose GitHub profile lists him as VP Engineering at Cognition, founder of Turborepo and formerly at Vercel and GitHub. Of 306 commits, 277 are his and 12 are from devin-ai-integration[bot]. No other contributor has more than two. The bus factor is one. Commit activity went from 64 on 20 September to 3 on each of the last two days.

That matters less than it would for a library. The adapters are on Hugging Face with SHA-256 checksums in a GitHub release, under Apache-2.0, and the inference path is a few hundred lines you could vendor. If the project stalls, what you lose is the training recipe and future checkpoints, not the ability to serve the ones you have.

Verdict

The sharpest criticism in the HN thread came from hbrn: these models only pay off when you "need fast response" and "can tolerate Jev's mediocrity compared to real frontier models", and "outside of fun demos these two rarely come together." Our numbers mostly agree. Kev-4B was 26 articles worse than Haiku on a four-way classification that a person gets right at a glance.

Where we disagree is the gating use case. A model that was never wrong above 0.6 across 140 real examples is a safe first filter, and vidarh's suggestion in the same thread (sorting agent bash commands into safe and unsafe before a bigger model looks) is exactly that shape. Adopt Kev if you already route large volumes to an LLM and can run a GPU or a 32 GB Mac server. Wait if you want one model that replaces the LLM call. What would change our mind is a Kev-4B that commits at 0.6 or higher on half of our archive, not a sixth, without losing that precision.

Your turn

If you run a decision model or a classifier in front of an LLM, at what confidence do you let it act without the big model, and what share of your traffic clears that bar? Our answer tonight was 0.6 and 16%. We would like to know whether that is typical or just our data.

The Laya run on the same 140 articles is worth reading next to this one: Laya committed more often and ran six times faster, and Kev was right more often. Which one you want depends on which of those you are paying for.

Share

Kev gained 4,310 GitHub stars in a week. On 140 of our own articles it lost to Haiku by 26, yet every answer it gave with 0.6+ confidence was right. #OpenSource #MachineLearning #LLM #Python

Never miss a ship

The best stuff that shipped this week, delivered every Thursday. Free, no spam. We read all the boring stuff so you get the fun parts.

Keep reading