Repo6 hours ago

Laya Added 24,000 Stars in Eight Days, Then Sorted Our Own Archive Worse Than Haiku

An open decision model went from 2,824 to 26,781 stars in under eight days. We ran it over 140 of our own articles: 64.3% right in 45 ms each, and the confidence score was the best part.

The WJS Desk

Sep 28, 2026 · 6 min read

Photo by BREAKS OUT on Pexels

The number that made us look

Our trending tracker first saw NandhaKishorM/laya on 20 September at 2,824 stars. By 02:12 UTC on 28 September it had 26,781, plus 2,339 forks. That is 23,957 stars in under eight days, and for two straight days it added more than 5,900 stars every 24 hours (4,981 at 04:00 UTC on the 21st, 17,300 at 04:00 UTC on the 23rd, measured by our own snapshots, not a badge).

It is slowing down hard. The last 22 hours before we wrote this added 824. The spike lines up with a Hacker News post on 19 September that sat at 1,358 points and 318 comments when we checked, and with the wider Jev craze: eight of the 25 repos in our tracker this week, Laya included, name Jev or Laya in their title or description.

The commit graph moves just as fast. The repo has 712 commits since its first on 18 September, 29 releases on PyPI, and 119 commits dated 28 September alone.

What it actually is

Laya is a decision model, not a chat model. You give it some text and a set of typed questions, and it answers all of them in one forward pass: a choice from labels you define, a score on a scale, or a noul yes/no probability. Nothing is generated, so there is nothing to parse. The README pitches it as an open, self-hostable counterpart to TypeSafe's hosted Jev API, and it ships an HTTP server that speaks Jev's wire protocol.

Under the hood there are three checkpoints on encoder models: 421M-parameter ModernBERT-large for English, 322M-parameter mmBERT-base for 100+ languages, and a fine-tuned typed-decisions variant. A Router detects the language and picks one. The core API is small:

from laya import Router

router = Router()
Q = {"category": {"type": "choice",
      "instructions": "What kind of article is this?",
      "criteria": {"news": "a launch, release or incident that just happened",
                   "tutorial": "a hands-on step-by-step guide",
                   "repo": "a spotlight on one trending GitHub repository",
                   "pulse": "what developers are arguing about online"}}}

r = router.predict(title + "\n\n" + excerpt, Q)
r["answers"]["category"]["choice"]             # "repo"
r["answers"]["category"]["answer_confidence"]  # 0.61

The code is not small. The laya/ package is 12,418 lines of Python, and the four files that do the actual inference and routing (agent.py, router.py, common.py, confidence.py) are 3,235. Direct dependencies are just five (torch, transformers, safetensors, huggingface_hub, numpy), but a clean install pulled 37 packages into a 714 MB virtualenv.

We tried it on something real

We needed a classification task with ground truth we trust, so we used our own archive. Every one of our 140 published articles has a human-assigned category. We fed Laya the title and excerpt of each, asked it for one of our four categories, and compared. Then we sent the identical prompt, with the same four definitions, to Claude Haiku as a single batched call through the Claude CLI.

The quickstart worked first time on an M4 Pro. uv pip install laya finished in 1.3 seconds, but only because torch was already in our uv cache, so do not expect that on a fresh machine. The first predict took 61 seconds, almost all of it downloading 1.6 GB of weights. After that the three-question quickstart ran at a 42 ms median over 20 calls on the Apple GPU.

140 of our articlesLaya 0.3.21 (english)Claude Haiku
Correct category90 (64.3%)108 (77.1%)
News correct17 of 3934 of 39
Tutorials correct35 of 5733 of 57
Time per article45 ms on Apple GPU, 88 ms on CPU191 s for one batched CLI call
Runs offline, no per-call costYesNo
Disk before the first answer714 MB venv plus 1.6 GB weightsNone

A caveat on that timing row: the Haiku number includes Claude CLI startup and is one call for all 140 articles, so it is not a per-request API latency and we would not quote it as one. The accuracy gap is the fair comparison. Accuracy was the same 90 of 140 on GPU and CPU.

Both models tripped on the same blurry line. Laya sent 16 of our tutorials to pulse, mostly the numbered parts of our multi-part courses. Haiku sent 24 tutorials to repo, because those courses are about a repo. Laya also filed 13 news pieces as pulse, which we suspect is our own fault: our news pieces say "we tested" a lot.

The interesting part is the calibration. Laya returns an answer_confidence with every answer, and on our data it meant something:

  • At 0.7 or higher: 27 answers, all 27 correct.
  • At 0.6 or higher: 52 answers, 49 correct.
  • Below 0.5: 68 answers, only 29 correct.

So we built the obvious cascade: take Laya's answer when it is at least 0.6 sure, send the rest to Haiku. It scored 110 of 140, two better than Haiku alone, and it skipped Haiku on 52 of 140 articles. That is the honest shape of it: Laya did not replace the LLM call, it removed 37% of them.

The hidden cost: you are now running a model. That means 1.6 GB of weights per checkpoint (the multilingual one is a separate download we did not take), a 61-second cold start, torch on every box that serves it, and a pinned revision, because this project cut 29 releases in ten days and a checkpoint change can move your accuracy without touching your code.

What it does not do

  • Many labels. Options share a fixed token budget (192 tokens on the English checkpoint). The maintainers' own table shows Banking77 at 0.425 for Laya against 0.870 for Jev, and they call it out themselves.
  • Long documents on the English checkpoint. It reads 512 tokens. Only the multilingual one stretches to 8,192, and the README reports accuracy getting patchy beyond roughly 4,000.
  • Generate anything. No explanations, no extraction. You get a label and probabilities.
  • Match Jev's confidence number. Laya's confidence field is computed differently, so thresholds tuned on Jev do not carry over. Gate on answer_confidence.
  • Beat a general LLM zero-shot on fuzzy categories. The README says fine-tuning is where accuracy jumps (0.362 to 0.766 on their typed-decisions set). We did not fine-tune, so our 64.3% is the zero-shot floor.

Who made it, and does that matter

The git log lists 110 authors in ten days, which mostly reflects a star spike turning into a pull-request spike. The weight sits with two people. Nandakishor, the author, has 270 of the 712 commits across two email addresses and 47 of the commits to agent.py and router.py. A contributor named Aashish has 117 commits under two names and the most substantial changes inside the package (41 commits of 20 or more lines). After those two, about 15 people have landed more than one commit of that size.

That is a better bus factor than most ten-day-old repos, but the 3,235-line inference core is too large to vendor casually, and the weights live on one organisation's Hugging Face account. If the project stalls, the Apache 2.0 licence lets you keep a pinned checkpoint and the package, which is the realistic fallback.

Verdict

Adopt it today if you already make LLM calls to route, triage or gate text into a handful of buckets, and you can put a confidence threshold in front of it. On our data, that threshold is where all the value is. Wait if you need 20+ labels, long documents, or zero-shot accuracy on categories that humans argue about, because a general model still wins those. What would change our minds: a fine-tuned checkpoint on our own archive clearing Haiku's 77.1%, which the maintainers' notebook runs on Kaggle's free T4 GPUs. We have not run it yet.

Your turn

If you have put a small classifier in front of an LLM call: what share of requests does it actually handle on its own at a threshold you trust? Ours was 37%, and we would like to know if that is low. For the tool that started this wave, we measured browser-use's version and found only part of the claimed saving came from where it said: the jev-ultrafast audit is the context for why Laya's star count exists at all.

Share

26,781 stars, most of them in one week. We ran this open decision model over our own archive: 64% right, 45 ms each, and every answer it was 70% sure of was correct. #OpenSource #MachineLearning #Python #LLM

Never miss a ship

The best stuff that shipped this week, delivered every Thursday. Free, no spam. We read all the boring stuff so you get the fun parts.

Keep reading