News5 hours ago

Grok 4.7 and Jev Shipped Six Days Apart and Share Zero Benchmarks

Grok 4.7 costs about 48 times more per input token than Jev, and its median time to first token is 54 seconds against Jev's 0.32. They are also not the same kind of model, so we worked out which job each one is actually for.

The WJS Desk

Sep 28, 2026 · 7 min read

Photo by Michael on Pexels

Two models shipped six days apart this month. TypeSafe launched Jev on 15 September, and xAI (which now signs its own site as SpaceXAI) released Grok 4.7 on 21 September. Readers asked us which one to use. The honest first answer: xAI published scores on ten benchmarks, TypeSafe published its own workflow evals, and not one benchmark appears on both lists. Neither vendor tested against the other.

So this is a paper comparison, and we want that clear up front: we have not run either model for this piece. We read both launch posts, both sets of docs and model pages, the Artificial Analysis measurements, one independent Jev test, and both Hacker News threads (609 points on Grok 4.7, 1,984 on Jev when we checked). Then we lined up the prices and the limits, and worked out which job each one is actually for.

Read this first: every number below is somebody else's. Vendor numbers are labelled as vendor numbers, independent ones name who measured them. Where the two sides measured different things, we say so rather than force them into one column.

What actually shipped

Grok 4.7 is a general-purpose reasoning LLM, API id grok-4.7. xAI's model docs list a 500,000-token context window, text and image input, text output, and a May 2026 knowledge cutoff. Pricing is $2 per million input tokens, $0.50 cached, and $6 output below 200,000 prompt tokens, doubling to $4, $1 and $12 above it. The launch post says a "fast variant" runs at twice the output speed for twice the price. It supports JSON-schema structured outputs, and xAI's docs say a response is "guaranteed to match your schema" when you stay inside the supported schema features.

Jev is not an LLM, and TypeSafe says so repeatedly. The current model is Jev 1.13, id jev-1.13.0, and both the jev-latest and jev-preview aliases point at it. You send it text "state" and typed questions, and it returns one of three things: a Choice from a list you defined, a Score on a rubric you wrote, or a Noul, a probability that a statement is true. It never writes text. Its models page lists $0.042 per million input tokens with output free, a 64,000-token budget per request (32,000 for the state plus the longest question), text-only input, 250,000 tokens per second and 1,200 requests per minute, with a warning that those limits "can change without notice".

TypeSafe's own coding-agent page is blunt about the overlap: Jev is "not a drop-in replacement for the LLM behind Claude Code, Cursor, opencode, Copilot, Muse Spark, Grok Bot, or similar tools." That sentence settles the framing. Nobody should pick between these for writing code. The real choice only exists for one job: a structured decision inside software, such as routing a request, classifying a document, screening a message, or picking the next action from a fixed list. Grok 4.7 can do that job through structured outputs. Jev only does that job.

The spec sheets side by side

Line itemGrok 4.7Jev 1.13
Input price, per 1M tokens$2.00 (xAI docs)$0.042 (TypeSafe docs)
Output price, per 1M tokens$6.00, reasoning tokens includedFree
Context per request500,000 tokens64,000 total, 32,000 for state
Input typesText and imagesText only, English strongest
What comes backText, code, tool calls, schema JSONA choice, a score, or a probability
Latency, as reported54.44 s median to first token (Artificial Analysis, xhigh)70 to 500 ms end to end (TypeSafe); 0.32 s median (Lindfors)
Can you sign up todayYes, public APINew signups paused since 22 September

Our arithmetic on the price line: $2 divided by $0.042 is 47.6, so Grok 4.7 costs about 48 times more per input token before it produces a single output or reasoning token. Tokenizers differ, so treat that as an order of magnitude, not a quote.

Why the benchmarks never meet

xAI's table is all generation and agent work: 46.3% on CursorBench 4.0 (against 51.8% for Fable 5.1 Max and 41.7% for GPT-5.6 Sol Max), 71.0% on DeepSWE v1.1 at high effort, and 37.6% on Terminal-Bench 4.0. Jev cannot attempt any of those, because every one of them requires writing code or text.

TypeSafe's launch post goes the other way. It reports Jev at "193.6x faster, 444.6x cheaper" on its workflow evals and 70 to 500 ms against "3 to 329 seconds" for frontier LLMs, comparing against GPT-5.6 Terra, GPT-6 Astra, Fable 5.1 and DeepSeek models. Grok is not on the list. TypeSafe also concedes the workflows "were made by individuals on our model capabilities team, so some bias could exist." The sharpest reply on the Jev Hacker News thread, from ramon156, called the latency comparison "apples-to-oranges unless the LLM baseline is doing comparable work."

So we checked each vendor's numbers against independent ones instead. Grok 4.7's shrink. Jev's latency claim holds up, on a small sample.

  • Grok 4.7 on Terminal-Bench 4.0: xAI reports 37.6%. Artificial Analysis measured 25.8% at xhigh and 24.7% at high. Same benchmark name, different harness and run, and a gap of nearly 12 points.
  • Grok 4.7 overall: Artificial Analysis scores it 46 on its Intelligence Index v4.3.2, 21st of 211 models. The Decoder reports Fable 5.1 and GPT-6 at 53 on the same index.
  • Grok 4.7 settings: xAI's headline table compares Grok 4.7 at xHigh with Grok 4.6 at High. HN commenter AM1010101 asked why the settings do not match; the post does not say.
  • Jev latency: an early-access test published at lindfors.no ran jev-1.13.0 on 18 September, classifying 24 Norwegian government hearing responses. The author measured a 0.32 s median and 1.3 s slowest request, against 26 s median for DeepSeek V4.1 Flash with reasoning on. That supports TypeSafe's range, on a sample of 24.
  • Jev accuracy: in the same test, Jev matched the reference label on 20 of 24 stances. DeepSeek with reasoning got 22 of 24. Cost per 1,000 documents was $0.22 for Jev and $3.08 for DeepSeek with reasoning.

The downsides, named

Grok 4.7 is slow and wordy for small decisions. Artificial Analysis lists a 54.44 s median time to first token at xhigh and 44.07 s at high, against a 3.79 s median for reasoning models in its price tier. Even the fastest 5% of its xhigh requests took about 10 s. It generated 240 million output tokens to complete the index, and HN commenter sourcecodeplz pointed out Grok 4.6 at xhigh needed 97 million for a score of 44. That verbosity is billed at $6 per million. GodelNumbering added that cached reads cost $0.50 per million on Grok against $0.25 on Fable, which matters for long agent loops. One hands-on report in the thread, from Saline9515, described it looping in thinking mode and ignoring AGENTS.md instructions. That is a single user, but it is the kind of failure a benchmark table never shows.

Jev is narrow by design, and currently hard to get. TypeSafe's own Jev 1.13 jaggedness page lists nine failure modes, including "does not count reliably", reading dates "as text, not as ordered quantities", accuracy that drops with large irrelevant state, and being steerable by injected instructions. Lindfors found that careful, precise instructions raised Jev's calibration error from 0.040 to 0.116, which is the opposite of what you would expect from an LLM. And access has flipped twice: TypeSafe dropped the waitlist on 20 September, then paused new signups on 22 September, saying existing accounts would keep working. We found no announcement that signups have reopened.

We have seen the calibration issue before. When we covered fast-jev-compaction, users who ran it against jev-latest reported that "Jev's ranking is correct and very stable; only the absolute calibration is off." If your code branches on a raw probability threshold, pin jev-1.13.0 rather than the alias, which TypeSafe's own docs recommend.

Our verdict

This is not a head-to-head, and anyone framing it as one is selling something. On paper, these are two layers of the same stack.

Pick Jev when the answer is one of a fixed set: routing, moderation, classification, reranking, picking an agent's next action. Your input also needs to be English text under 32,000 tokens, and you need to already hold an account. On every number published so far it is roughly 48 times cheaper per input token, bills nothing for output, and answers in well under a second where Grok 4.7 at high effort takes tens of seconds. Pin the version and set thresholds from your own data, not from the probability alone.

Pick Grok 4.7 when the output has to be words, code, or tool calls, when you need more than 64,000 tokens of context, when the input includes images, or when the decision needs multi-step reasoning that Jev's jaggedness page says it handles badly. Just know what you are buying. Independent numbers put it 7 points behind Fable 5.1 and GPT-6 on the Artificial Analysis index and well behind them on agentic terminal work, so its case is price, not the lead.

Where data is genuinely missing: nobody has published Grok 4.7 and Jev on the same classification set, and Artificial Analysis has not measured Grok 4.7 at low effort, which could close some of the latency gap. Until someone runs that test, the safe bet is Jev in front making the cheap decisions, and an LLM like Grok 4.7 behind it for the calls that need one.

Your turn

If you route or classify with an LLM today, what is your median latency per decision, and at what number would you move that step to a decision model? For why Jev spawned an entire ecosystem of wrappers in its first week, and how much of one popular wrapper's claimed speedup held up when we measured it, read our jev-ultrafast audit. If you want to self-host rather than wait for signups, our Laya spotlight covers the open clone that speaks Jev's wire protocol.

Share

Grok 4.7 costs about 48x more per input token than Jev and takes 54s to its first token. Jev cannot write a sentence. We lined up both spec sheets. #Grok #xAI #LLM #AI

Never miss a ship

The best stuff that shipped this week, delivered every Thursday. Free, no spam. We read all the boring stuff so you get the fun parts.

Keep reading