BTC $84,012.68 -0.89%
ETH $2,697.56 +0.03%
BNB $769.85 -1.11%
XRP $1.51 -1.67%
SOL $119.96 -2.59%
TRX $0.3363 +0.59%
DOGE $0.0949 -2.82%
ADA $0.2494 -3.00%
BCH $310.81 -7.49%
LINK $15.25 +7.86%
HYPE $88.75 -3.54%
AAVE $149.22 -4.21%
SUI $1.15 -8.14%
XLM $0.2305 +5.99%
ZEC $1,512.57 -6.17%
AAPL $339.28 -0.29%
AMZN $246.27 -1.52%
GOOGL $342.46 -0.42%
MSFT $511.34 -1.14%
META $720.84 -3.87%
NVDA $229.71 +1.95%
TSLA $358.34 -4.00%
SNDK $1,712.55 -3.96%
INTC $115.42 -8.11%
SPCX $147.33 -1.02%
MU $1,053.67 -3.95%
AMD $607.16 -4.16%
BTC $84,012.68 -0.89%
ETH $2,697.56 +0.03%
BNB $769.85 -1.11%
XRP $1.51 -1.67%
SOL $119.96 -2.59%
TRX $0.3363 +0.59%
DOGE $0.0949 -2.82%
ADA $0.2494 -3.00%
BCH $310.81 -7.49%
LINK $15.25 +7.86%
HYPE $88.75 -3.54%
AAVE $149.22 -4.21%
SUI $1.15 -8.14%
XLM $0.2305 +5.99%
ZEC $1,512.57 -6.17%
AAPL $339.28 -0.29%
AMZN $246.27 -1.52%
GOOGL $342.46 -0.42%
MSFT $511.34 -1.14%
META $720.84 -3.87%
NVDA $229.71 +1.95%
TSLA $358.34 -4.00%
SNDK $1,712.55 -3.96%
INTC $115.42 -8.11%
SPCX $147.33 -1.02%
MU $1,053.67 -3.95%
AMD $607.16 -4.16%

Delving into the fundamentals to see if Jev can wear the crown of "paradigm innovation."

Core Viewpoint
Summary: General judgment requires the simultaneous use of multiple abilities; one cannot assume it is easier than generating an answer just because it ultimately outputs a single probability.
Tencent Technology
2026-09-27 00:04:32
General judgment requires the simultaneous use of multiple abilities; one cannot assume it is easier than generating an answer just because it ultimately outputs a single probability.

Author: Boyang

Editor: Xu Qingyang

In the tech circle, the hype cycles are always astonishingly similar.

In September 2026, the entire Silicon Valley and open-source community were fervently excited about a model named Jev.

It claims to be a "System One" model, self-proclaimed to bring disruptive changes.

In the past three to four years, we have become accustomed to large language models (LLMs) that, like verbose essayists, always generate hundreds of meaningless thought tokens before slowly spitting out a JSON-formatted conclusion in response to a simple "yes" or "no."

But Jev is different; it promises to give you judgments directly, provide probabilities, and do so at an astonishing speed.

Developers have gone crazy for this "fast, accurate, and ruthless" interface.

After all, when building a complex AI agent, what you often need to ask is: "Is this memory relevant?" "Which tool should this request be assigned to?" "Does this incident require manual review?"

In business, everyone has spent too many tokens and delays to make a general large model generate text, explanations, and structured code for every small issue. By addressing real needs, Jev hits an extremely painful pain point.

But does it really have enough disruptive power?

When we skip the "text generation" process, how much of the judgment ability left in this black box comes from the original large language model's foundation, and how much comes from what the TypeSafe team calls specialized training?

To answer this question, we can peel back the layers of Jev's packaging and discuss it from several angles.

01

Predictive judgment models predate GPT

The direction represented by Jev has a very long history.

Delving into the fundamentals to see if Jev can wear the crown of

As early as 2018, Google launched BERT. Unlike the GPT-like models we are familiar with today, it adopted an Encoder-only architecture that allows for bidirectional information exchange, training the model to fill in the blanks.

Although later, Decoder-only models (like ChatGPT) became mainstream for training continuation, BERT still has its advantages in clear classification tasks.

With the ability to see all information, BERT can combine various positions in the text with context to form complete feature representations, then connect to a simple classification layer, and after training with labeled emails, it can quickly make judgments.

Modern LLMs often also play the role of classifiers.

In 2019, OpenAI began experimenting with getting models to learn human preferences for text in the paper "Fine-Tuning Language Models from Human Preferences."

By 2020, in research on text summarization, this technical pipeline became clearer, where human annotators first compared two summaries and then trained a reward model to learn which one humans preferred.

In this process, human judgment was successfully transformed into a scoring function. The reward model itself does not need to write lengthy comments; it only needs to read the question and answer, then directly provide a score through the output layer of a neural network.

This is actually one of the most fundamental patterns of model learning today, which is what Jev strongly criticizes as RLHF.

Therefore, "inheriting the understanding ability of language models but not generating text" is not a unique genius idea of Jev.

By 2023, in the well-known paper "Let's Verify Step by Step," this supervision was further refined and applied to every step of the model. Reward models, validators, and evaluation metrics developed from different tasks gradually formed several indispensable judgment tools in the AI industry.

By 2025 to 2026, related research was still advancing. In 2025, Galileo released Luna-2 and demonstrated in subsequent papers how to train small language models into "single-token classifiers," directly reading the probability of target categories through a single forward computation. Meanwhile, Skywork-Reward-V2 launched a series of reward models ranging from 0.6B to 8B, optimizing the direct scoring route for LLMs.

From the perspective of algorithm implementation, this is not inherently difficult. Making some simple changes to LLMs to calculate candidate scores directly from internal hidden representations, rather than spewing out a bunch of words, has many methods.

In the subsequent replication experiments, we can see more than three.

So what exactly has Jev brought that is different in this wave?

From the current architecture decryption and official explanations, Jev's core differentiation is reflected in two dimensions.

Delving into the fundamentals to see if Jev can wear the crown of

First, it is a more general change, thanks to its brand new post-training method.

Past scorers were basically trained for single tasks. Jev is no longer limited to a specific scoring type but attempts to provide an extremely general probability interface.

Moreover, this generality is about factual probabilities, not human preferences.

This time, the TypeSafe team placed "probability calibration" at the absolute center of the training objective, referring to this method as RLCD (Reinforcement Learning for Calibrated Decisions). Traditional preference reward models (RLHF) look for probabilities of human preferences rather than probabilities of real-world events.

Second, it is the extreme exploitation of parallel computing in architecture.

Jev changed the way information flows through the model. Past scoring models had little modification for parallel answering because the main goal was to train dense reward models, scoring each token that came out.

But this time, Jev can answer up to 250 questions on the same material. As long as there are no sequential dependencies between these questions, they can be executed in large-scale parallel batches.

How is this achieved?

Although Jev itself has not disclosed its architecture specifically, based on its actual performance and numerous attempts to replicate Jev, we can roughly describe its outline.

02

Piecing together Jev's original form from tests and replications

Let's first look at what TypeSafe has currently disclosed.

Outside the black box, TypeSafe has clearly published the interface. You can provide two things to this interface:

State: Equivalent to the original text of a long reading comprehension (for example, a lengthy customer complaint record or system log).

Questions: For this original text, you can simultaneously pose several questions. The questions do not interfere with each other. After the system receives these independent answers, it is up to the coder (you) to decide how the next business logic proceeds.

To standardize the questions, TypeSafe has narrowed the questioning methods down to three primitives. They include:

Noul (Yes/No questions): Ask "Is it?" and directly return a probability between 0 and 1 (for example: Is this urgent? Return 0.95).

Choice (Multiple-choice questions): Ask "Which one to choose?" provide several options, and return their respective probability distributions (for example: Forward to the tech department 0.8, forward to the finance department 0.2).

Score (Scoring questions): Ask "To what extent?" and return the probabilities of various levels and the final weighted score (for example: Customer anger index 4.5).

If the goal is to let the program make common judgments, these three primitives cover a large portion of output forms, and almost all judgments can be transformed into these three types of questions.

Finally, the program receives the numbers and takes action according to its own rules.

The official promise is that Jev will only read the State once. Subsequently, all questions will make parallel and independent judgments on this state within the same request.

Independence means that multiple questions cannot reference each other's answers. If your second question must rely on the result of the first question, you can only send requests in two batches.

Therefore, TypeSafe encourages a usage called "Speculative fan-out": for example, when handling a customer complaint, even if it turns out that this is not a system fault, the program can initially ask "Is it a fault?", "How serious is the fault?", "Who should it be forwarded to?" All answers can be calculated in parallel at once, and then downstream code logic can discard the useless results.

Delving into the fundamentals to see if Jev can wear the crown of

Currently, this is all the main information we know that has been officially disclosed.

After Jev gained attention, Archer Hume conducted a series of black box tests, from which we gained many insights into how Jev processes state text. Meanwhile, open-source replication projects like Kev, NanoJev, and minojev have sprung up like mushrooms.

By comparing which replication projects' test feedback is closer to the real Jev, we can reverse-hypothesize their true internal architecture.

Although it cannot be said that these replications 100% restore the essence of Jev, there are huge black holes left between the official interface documents, such as how the original text shares the state? How are ordinary candidate options formed into feature representations? At what step do they influence each other?

Through this lens, we can also glimpse a rough outline.

Next, we will use a specific request to retrace the "internal digestion" process of this black box large model.

Delving into the fundamentals to see if Jev can wear the crown of

First, we need to clarify the information structure that Jev actually processes. A complete handling includes three parts: the State (for example, that lengthy complaint text) as shared background, several independently posed Questions, and the corresponding Options for each of these questions.

First stop: processing common information

If each question has to read through tens of thousands of words of complaints from start to finish, the more questions there are, the more redundant computational power is wasted.

So the best method is for all questions to be able to read the question only once and then use it.

To verify whether Jev really achieved "reading only once," researcher Archer Hume reviewed the API's billing statements and latency data: when submitting the simplest yes/no question, the billing showed 268 input tokens; when increasing to two questions, it became 276. The extra amount was merely the word count overhead of the new questions themselves, and the system did not charge for the repeated reading of the common State material.

At the same time, before the number of questions increased to nearly a hundred, the server's response time was almost a flat line.

Although the commercial billing rules and batch processing masked the true computational records of the graphics card, this phenomenally aligns with the official claim that "shared materials are read only once, and each question is calculated in batches."

For this information architecture, the restoration of the open-source project Kev is currently the clearest.

Delving into the fundamentals to see if Jev can wear the crown of

Kev first processes the status all at once, freezing the intermediate results obtained from calculations in the Kv Cache. Next, these 50 questions share this Kv Cache for further calculations, so there’s no need to read each question again.

Second stop: Question Splitting

But another key question is, are these batch-calculated questions really isolated from each other as the official claims?

Archer Hume designed a clever "code experiment" for this. He inserted the phrase "the code is ZEBRA-7741" into question A, and then in the options for question B, he had the model select "the code mentioned in another question." The results showed that Jev's probability of giving the correct code was 0.00. However, if this code was removed from question A and placed into the shared State text, the probability of question B giving the correct answer instantly soared to over 0.90.

This constitutes strong evidence that there is strict physical isolation between the questions: shared materials are visible to all questions, but adjacent questions absolutely cannot "sneak a peek" at each other.

To ensure that the questions can be separated, Kev used two methods.

Delving into the fundamentals to see if Jev can wear the crown of

The first method is the attention mask. With it, when the system calculates [frozen original complaint] + [question 1] + [question 2] together, as long as the model is processing question 1, the masking mechanism will forcibly turn the area of question 2 into a state with a value of 0, forcing it to "pay attention" only to the shared original complaint and itself while answering.

The second method is independent branch reuse. When using bases with cyclical features or specific architectures (such as the Qwen3.5 mentioned in the text), the attention mask cannot be separated. Therefore, when using these models, once the model reads the original complaint, it will take this frozen memory as a starting point and directly split into 50 parallel highways (independent branches).

Since each branch perfectly inherits the already processed complaint memory from the main road, they also enjoy the benefit of not having to reread the original text.

Third stop: Calculating Options

Now that the state is shared and the questions have been split, how does the model calculate and process those candidate options within a split multiple-choice question?

At this point, the model has several candidate options in front of it, such as "Finance" and "Technology." The most traditional approach is linear head (Linear Head) plus Softmax. This is the pattern that the Zefan Open-Jev version attempted to replicate Jev.

You can equate it entirely to an absolutely closed "black box blind review." The contestant "Finance" performs in the black box, and based on a rigid scoring guide (this is the function of the linear head), you give them an absolute score of 80, then "Technology" enters the black box, and you give them a score of 90. These two never meet, and you never compare them. Finally, you use the Softmax mathematical formula, which specifically calculates percentages, to convert the absolute scores of 80 and 90 into winning probabilities.

In this hypothetical process, if we insert a completely illogical distraction, such as "bad weather," it at most serves as cannon fodder in the denominator, causing everyone's percentage to shrink a little. However, since the scores of 80 for Finance and 90 for Technology are already written in pen on paper, their "relative odds" can absolutely not be shaken by the addition of a distraction.

Delving into the fundamentals to see if Jev can wear the crown of

However, tester Archer Hume's tests proved that the addition of new options does indeed affect the score difference of the model. He forcibly inserted the distraction "bad weather" into a set of normal options and found in ten randomly arranged tests that it indeed changed the relative odds of Finance and Technology. This directly sentenced the "black box blind review" model to death.

Since it is not a blind review, it indicates that the contestants must have "seen each other" and generated a chemical reaction before the judges gave the final scores. The open-source community has provided two designs to achieve this chemical reaction.

Delving into the fundamentals to see if Jev can wear the crown of

The first scheme is the "Pointer Head" model designed by the Kev project. You can think of it as a "group interview." The model no longer locks contestants in a black box but lines up "Finance, Technology, Bad Weather," allowing the judges to see them all at once.

When the judges see the contestant at the end of the line, they have already formed an "overall context" about this interview in their minds. Then, standing at the last position, the judges point back to the previous contestants one by one, scoring based on their overall impression at that moment; this action is called the pointer head. When the judges look back to point and score after seeing everyone, their mindset and reference frame have changed, and the scores given to Finance and Technology naturally fluctuate accordingly.

The second scheme is "Internal Discussion of the Judges," designed by projects like NanoJev, which is the "Inter-candidate Attention Module" model. This time, the model is neither purely blind review nor purely group interview. It first lets "Finance" and "Technology" perform separately, condensing their performances into a long string of high-dimensional numerical comments, a term known as feature vector.

At this point, the two comment cards are still isolated. But then a separate small model will throw these two comment cards, along with the later inserted "bad weather" comment card, into a conference room called the "attention module."

Delving into the fundamentals to see if Jev can wear the crown of

In this conference room, these groups of numbers representing the contestants (feature vectors) will be compared and weighed against each other. Originally, Finance and Technology were difficult to distinguish, but suddenly the addition of a "bad weather" card disrupts and reorganizes the focus and comparative weight of the entire judges' discussion.

Delving into the fundamentals to see if Jev can wear the crown of

After this internal meeting of mutual undermining, the final scores given will naturally no longer resemble the initial two-person scenario.

Why does Jev go to such lengths to make these options "undermine each other" in the underlying code? This is not merely for show, but because in real complex business scenarios, the correct answer is often not absolute, but rather "compared." The options themselves are hidden clues to solving the problem.

For example, if asked where the Eiffel Tower is, the candidate options are A. Europe B. France C. Paris. When the model sees these three options simultaneously, these options themselves become hidden clues to deciphering the intent of the question. This question is not testing approximate location but is testing the highest precision of geographical location.

Fourth stop: Providing Results

The final step of the process is to derive a specific score.

According to TypeSafe's official statement, Jev will directly return a probability number, never generating text word by word.

Archer Hume's external detection also confirmed this; when he increased the candidate options from two to two hundred, the response text returned by the API, although much longer, did not proportionally extend the server's processing time.

This indicates that the large model indeed skips the most time-consuming autoregressive generation (which is predicting the next word like ChatGPT).

So, where does this final probability number come from? The technical implementation is actually varied, mainly relying on the methods used in the previous step of calculating options.

The approach of openjev/openjev, which does not specially handle option interactions, is to directly read the original scores (Logits) of the specified candidate letter tokens from the large model's internal vocabulary at the "answering position" where the model is supposed to generate the first character.

With the pointer head, Kev directly outputs comparative scores based on the global perspective provided by the pointer head. The minoejev, which has a shared module, relies on that manually added shared scoring module to produce results.

Regardless of which extraction method is used, as long as the process is designed properly, the large model can completely abandon verbose text generation and directly extract precise mathematical probabilities at the end of the computation, completing a system-level automatic decision.

At this point, we can clearly outline Jev's architectural model based on the evaluation and restoration.

Delving into the fundamentals to see if Jev can wear the crown of

This architecture itself is not complicated, and it is difficult to find parts that can be called paradigm-shifting. Although it is indeed an excellent engineering optimization for specific scenarios, that is all there is to it.

03

The guarantee of accuracy is post-training

The architecture only guarantees speed; how is Jev's high accuracy achieved?

It relies on what TypeSafe calls RLCD (Reinforcement Learning for Calibrated Decisions) post-training method. Since RLCD is a non-public black box, we can only look at the potential topics and algorithm implementations from the attempts in the open-source community.

The birth of a question for RLCD training

The simplest way is direct synthesis, such as the method provided by the Hmm version replica.

Researchers had DeepSeek V4.1 list over a hundred work scenarios in the code, including refund processing, fault troubleshooting, relevance retrieval, email diversion, etc.

Each time, select a scenario and pair it with material forms and question requirements. For example,

Using a long message with irrelevant details, write several refund cases.

Add a case that is easily misled by keywords.

Each case is accompanied by four to five multiple-choice, true/false, or rating questions.

DeepSeek V4.1 Flash then generates the entire set of materials, questions, candidate options, judgment criteria, and answers.

After that, testing whether this question can be used relies on hiding the original answers and letting DeepSeek V4.1 Flash answer again, calling it three times by default, requiring each time to provide option probabilities. Questions that allow at least two valid responses can be used.

Kev's method is to unify existing news classification, comment sentiment, text implications, and other datasets into a judgment format of "material + question + candidate answer" according to rules.

The model then generates rules and facts, calculates answers, and writes them down.

On September 24, Kev-4B also attempted to construct questions using real work environments. They collected 5,219 real consumer financial complaints and created questions around "what products are involved, what the main issues are." After the questions were generated, only when the judgments of two different teacher models were consistent with the original labels filled out by consumers would the labels be retained.

Delving into the fundamentals to see if Jev can wear the crown of

To make this training more effective, the restored test questions will also specifically produce some paired trap questions.

For example, two questions with the same rules but changing one key name (for instance, changing the signer from the authorized Mira to the unauthorized Noah) will directly flip the answer. Such questions can effectively prevent the model from hardcoding answers through Reward Hacking, forcing it to honestly learn the deep representations and connections between questions and answer options.

Delving into the fundamentals to see if Jev can wear the crown of

Is training method, LoRA plus distillation enough?

Currently, almost all reproductions using post-training methods employ the "LoRA + teacher distillation" approach to increase the probability of correct options.

Delving into the fundamentals to see if Jev can wear the crown of

Taking Winnow as an example, it chose to perform LoRA fine-tuning on the Gemma 4 12B instruction model. The role of LoRA is to retain the original weights of the base model while training a small portion of correction parameters involved in the computation, reducing the cost of modifying the model.

Winnow simultaneously uses two types of supervision. One provides standard answers, requiring the model to increase the probability of the correct option. The other provides the teacher's probability distribution across all options, guiding the student to align with this distribution. Winnow only adopts the second type of supervision when the first choice selected by the teacher matches the standard answer.

Both are trained using cross-entropy to calculate training errors. The role of cross-entropy here is to check how much the student misallocated probabilities. The lower the probability assigned to the correct option when only the standard answer is present, the greater the penalty.

When using the teacher distribution, training will push the student to imitate the teacher's allocation across options. For example, if the teacher allocates 80%, 15%, and 5% to options A, B, and C respectively, the student will be required to learn the distinctions between these three options.

Generally speaking, LoRA, distillation, and cross-entropy are sufficient for tasks that already have prepared questions, answers, and reference distributions (such as this probability prediction task) because they are only learning a probability.

However, distillation actually learns the teacher's probability, not the real probability as mentioned by RCLD. How the gap from distillation crosses over through facts/teacher ratios is currently unclear in reproductions.

Theoretically, to achieve Jev's claims, it can only rely on larger data volumes and more sample-efficient learning methods.

Of course, if such a small judgment model can achieve the judgment accuracy of a large model, it would still be very useful even if it hasn't solved this issue.

Calibration may still be a secret weapon.

In addition, Jev has another secret weapon, which is its claimed calibration ability.

The number of questions a model answers correctly is separate from whether it accurately expresses its confidence. A model might only answer 70% correctly but report a 90% confidence level.

To verify whether Jev is blindly confident, Archer Hume conducted a "lie detection experiment."

He first fed Jev 1,200 MMLU (Massive Multi-task Language Understanding) test questions, then divided all selected answers according to the probabilities reported by the system into ten levels. After complex weighted calculations, Jev's calibration error was only 0.031.

In 30 simple three-digit multiplication questions, Jev answered 86.7% correctly, while its reported average confidence was 83%, which was very consistent. When the questions changed to more difficult "two-step application problems," its accuracy plummeted to 32%, the key point is that its reported average confidence also obediently dropped to 30%.

In this regard, post-training can indeed improve probability performance, but often it makes the model more blindly confident.

To restore Jev's relatively accurate calibration ability, the replica version used some methods. For example, Kev explicitly requires the model to reduce certainty in the absence of evidence. It adds samples where key evidence has been removed in the training set, then lowers their probabilities in the answers, so the model is penalized for baselessly concentrating probabilities on a certain option.

However, more reproductions merely perform a superficial temperature calibration.

Delving into the fundamentals to see if Jev can wear the crown of

Engineers will present a small batch of test questions that the trained model has not seen. If they find the model overly confident, the system will calculate how much the overall confidence should be adjusted based on this batch of questions.

Once the knob is adjusted, the model will be forced to wear a humility filter when outputting probabilities in the future, resulting in a more moderated probability.

This will not change the ranking of options for the same question, so the highest probability answer remains unchanged. However, the probability threshold and scoring based on probabilities may change.

This method is still quite effective. For example, after temperature calibration, the calibration error of Kev-9B dropped from about 10.6 percentage points to 4.2 percentage points, while the number of correctly answered questions remained unchanged.

Saying it is a superficial fix is because this actually comes from lowering overall confidence rather than from more accurately distinguishing whether it should be confident.

If Jev has indeed achieved an effective improvement in calibration rate, then they may indeed have some unique skills here.

04

What is the extent of Jev's applicable boundaries?

There is no doubt that Jev has practical application significance. It restores tasks that originally only required quick decision-making back to quick decision-making itself.

In the entire process of Agent, there are many explicit requirements for request classification and routing, retrieval result sorting, and detailed checks. These are all areas where Jev can play a role.

For example, customer service routing, product categorization, feedback analysis, and data labeling are all common high-frequency scenarios in our daily tasks.

But how large its applicable boundaries are determines whether it truly has the paradigm value it claims.

From the current benchmarks, it does indeed have a certain degree of generality. The Nimble team tested fact-checking, intent routing, semantic entailment, content review, medical Q&A, and other tasks using 13 groups of a total of 3,880 public data, with Jev achieving an average accuracy of 76.0% across datasets.

This indicates that Jev can indeed handle various judgment tasks.

However, true generality faces two hurdles.

The first is difficult tasks; if it can only make simple judgments, its applicable range is very limited.

First, we need to define what constitutes difficulty in judgment. Generally, we believe that the more steps, conditions, and detailed understanding required, the more difficult the problem.

Identifying "the user wants a refund" only requires understanding the literal meaning, but judging "should this refund be approved" requires checking dates, calculating deadlines, and comparing the priority of terms, which is clearly more difficult. If a judgment requires three steps of reasoning to reach a conclusion, then the third step depends on the result of the second step, which is multi-step dependency. If the policy is misidentified in the first step, even if the subsequent calculations are completely correct, the final judgment will still be wrong.

In the difficult question tests of JevBench, JevBench defined several factors related to the difficulty of judgment questions, including multi-condition judgments, continuous evidence searches, and comparisons of dates and numbers among three options. In this batch of questions, Jev's multi-step search accuracy was 85.7%, long policy judgment was 60.5%, and time and number judgment was only 26.7%, all far inferior to Flash-level models. In total, Jev answered 74.1% of difficult questions correctly, while DeepSeek V4.1 Flash achieved 95.0%, roughly on par with human experts.

Delving into the fundamentals to see if Jev can wear the crown of

Although the median response time for the corresponding tests was about 0.67 seconds and 3.15 seconds, Jev's speed was nearly four times faster. However, such a large accuracy gap makes it clear what choices to make for judgments facing complex tasks (such as stock trading, which everyone hopes for).

Moreover, the multi-step search accuracy mentioned here is not multi-step reasoning, but mainly involves searching. In Archer Hume's tests, Jev achieved a high accuracy of 86.7% in simple multiplication, but when switched to two-step application problems, the accuracy suddenly dropped to 32%.

Briantrust conducted a more detailed evaluation, which also highlighted Jev's shortcomings on difficult questions. It compared Jev with the GPT 5.6 Luna model, asking the model to select the correct one from two candidate answers, ultimately comparing 616 valid question pairs. The knowledge question gap between the two models was only 1.6 percentage points, while the gap in mathematics and reasoning expanded to 19.3 percentage points, and in coding, it reached 20 percentage points.

Therefore, regarding the hurdle of difficult judgments, Jev can hardly claim to have overcome it.

The other hurdle is generalization.

If Jev can truly generalize, it should be able to achieve better generalization, enhancing the accuracy of judgments on other questions based on what it has learned.

Otherwise, it is merely a "specialized judgment model" with a slightly broader applicable range, not significantly different from previous judgment models.

Existing multiple direct testing evidence can prove that Jev does indeed possess a certain degree of generalization ability, but this ability is extremely unstable within the domain. Once it steps out of familiar scenarios, its performance still faces significant risks.

First, under completely identical tasks and facts, Jev's judgments are easily influenced by the way information is presented. In the OpenProse experiment, without changing any facts, questions, or computational loads, when key relationships were concentrated at the front of the text, Jev's accuracy was 80.5%, but when these relationships were moved to the middle, the accuracy plummeted to 40.9%.

This indicates that even without changing business domains, Jev's robustness in representation is severely lacking. Its judgment ability highly depends on how external programs feed it data.

Secondly, when faced with new questions under the same evaluation system, Jev's performance also fluctuates significantly. In the new version v1.4 of JevBench, which added 308 closed difficult questions, Jev's accuracy dropped from 86.6% on public questions to 36.7%. In contrast, the DeepSeek V4.1 Flash using a thinking mode maintained at 94.8%.

The high scores Jev previously achieved on public questions do not automatically carry over to unknown complex tests.

When testing further pushes towards real business migration, Jev's performance is equally mixed. A positive example comes from the prompt injection detection of Agent Journal, where Jev achieved an accuracy of 83.65% on the original test set, and when directly migrated to another test set containing over two thousand external data points, the accuracy even improved to 95.58%, far exceeding traditional baseline models.

This indicates that it can indeed adapt to changes in data sources for certain specific tasks.

However, in more complex business logic, this generalization often fails. Scarif Labs once used it to determine whether software updates were safe. When moving the model from one software ecosystem to another, Jev's ability to distinguish between safe and dangerous updates (AUROC) dropped from 0.851 to 0.605 (close to the random guessing rate of 0.5).

Delving into the fundamentals to see if Jev can wear the crown of

Therefore, we can at least say that Jev's generalization is currently quite limited and questionable. It is still far from the standard of universal generalization.

Since Jev has not overcome these two hurdles of universality, where should we use it now?

A more suitable positioning is to place Jev in scenarios where the standards are clear, evidence is concentrated, and the judgment results can be verified or corrected.

Taking refunds as an example. Determining whether a user has expressed a desire for a refund, whether the complaint involves logistics or product quality, and whether additional proof is needed are all semantic judgments that can be verified individually. These questions arise frequently, usually do not require generating an explanation, and do not need the main model to engage in reasoning every time. Jev can take on preliminary triage here.

However, approving a refund changes the situation. Whether it exceeds the deadline requires code to verify dates, refund amounts need to be calculated according to rules, and in cases of conflicting terms or exceptions, further review is necessary. TypeSafe itself also suggests delegating mathematical operations and date comparisons to code and minimizing multi-layer dependencies in judgments.

This division of labor allows its speed to be utilized in appropriate places.

As for uses that rely more on precision, such as reward models, the judgments often involve those seemingly reasonable but actually incorrect answers. The model also needs to discern subtle differences and resist interference from expression styles.

For tasks where the rules are difficult to define, such as those involving relatively subjective evaluations, it is advisable to first test Jev's accuracy before deployment.

At this point, Jev can serve as a candidate solution, but users still need task-specific evaluations.

Delving into the fundamentals to see if Jev can wear the crown of

In addition, there is one account that cannot be overlooked. Jev's quick response does not mean that the entire Agent becomes faster after adding it.

If it can handle a batch of simple requests in advance, preventing these requests from calling the main model, then the speed and cost advantages can be realized.

However, if Jev is called first every time and then the same main model is still called, the additional judgments must save enough subsequent work to offset its own time and costs.

In a set of Agent memory retrieval experiments on GitHub, the system introduced Jev to determine whether the retrieved information was useful. To avoid Jev mistakenly discarding useful information, developers had to repeatedly adjust the judgment criteria.

Although ultimately all 20 conventional cases passed, the cost was clear, with the system's total latency rising from 649 milliseconds to 1087 milliseconds, and the cost per thousand calls more than doubled.

If the main model itself has the ability to filter answers from complex materials, adding Jev as a judgment layer not only adds waiting time but also carries the risk of missing key evidence due to misjudgment.

05

How true is Jev's disruptive potential?

In the end, the core question Jev leaves us with is, what exactly is it learning?

Logically, a model used for judgment learns representations through new facts to make effective judgments.

This is one of the most challenging representations.

Universal judgment requires the simultaneous invocation of multiple capabilities; it cannot be assumed to be easier than generating answers just because it outputs a single probability.

Take refunds as an example. Seeing "I want a refund," recognizing the desire for a refund primarily relies on language understanding. But asking "According to this policy, should the refund be approved?" requires the model to understand the policy, map facts like purchase dates and product conditions to specific terms, and then handle exceptions.

This is essentially the foundational capability of the language model behind Jev.

The final question, "Does the refund help retain this customer?" requires predicting behavioral outcomes. This is the part the model needs to be trained on.

It needs to learn which facts influence the outcome, under what conditions they influence, and whether these relationships still hold in different scenarios.

All three questions can output a "yes probability," but the knowledge and calculations required behind them are different.

To bridge the gap from "predicting how others will judge" to "reliably predicting the consequences of actions," how much data and training do we actually need?

At least for humans, comprehensive decision-making is one of the hardest things to learn. It is essentially as complex as the taste and comprehensive utilitarian calculations mentioned in our previous articles.

Therefore, I am very skeptical that the learning methods revealed in this article can support such complex representations.

Of course, we can view Jev as a highly ingenious engineering tool.

In business pipelines where rules are clear and materials are complete, it can significantly reduce waiting time and computational costs, and its value is undeniable.

However, before proving that it has indeed learned some universal judgment rules, it is premature to crown it with the title of "paradigm innovation."

Join ChainCatcher Official
Telegram Feed: @chaincatcher
X (Twitter): @ChainCatcher_
warnning Risk warning
app_icon
ChainCatcher Building the Web3 world with innovations.