Jev and the Return of Fuzziness

You are absolutely right!

I have lost count of how many times Claude had said that to me at work (in fact there used to be a website just to track it). The pandering issue of AI has never been this real, not just with Claude, but also Google AI, ChatGPT.., you name it - they could be excessively encouraging and asserting in response to your opinions, even if they’re far from truth. You’d hear constant compliments and validations like “that idea was perfect / a stroke of genius” etc.

That’s partially due to the fact that under the hood of these AIs, they typically have a rewarding function that’s by nature a pleasing personality, much like Mr. Meeseeks in Rick and Morty, they’ll stop at nothing until the “problem” is resolved - “how” may not be that important. To quote Mr. Meeseeks:

We Meeseeks are not born into this world fumbling for meaning, Jerry! We are created to serve a singular purpose for which we will go to any lengths to fulfill! Existence is pain to a Meeseeks, Jerry! And we will do anything to alleviate that pain!

mrmeeseeks

They can obviously do goodness with this level of conviction and can-do perseverance, and do so many of them, so much so that it’s getting impossible to keep up - there’s always something new every week. I always joke to my friends that, at this AI age, a little procrastination probably does more good than bad - who knows, if you defer learning something by a couple of days, you may not have the need to learn it at all as it could well be obsolete.

In fact, today I saw in the news that OpenAI dropped 722 math papers written by internal reasoning model, claiming that it’d solved some of the hardest cold math problems in the past century. As a math major myself, I started to feel sorry for the career mathematicians - how are they supposed to compete with the most powerful AIs with pretty much infinite computing resources? More importantly, they’d now share what us fellow software engineers suffer - the review fatigue. At the end of the day, these AI-generated deliverables, however fancy and formal they look like, would still have to be peer-reviewed (is it even peer tho) and verified by humans. Or maybe this really is the end of mathematics as a human-led science (perhaps similar to SaaSocalypse, a Mathocalypse?)1?

Although, these AI models don’t really have the most convincing track record. In fact, just not so long ago, an Ontario man named Allen Brooks sued OpenAI after alleging its ChatGPT over-validated his ideas, agreed with him unconditionally, and eventually caused him distress and loss of career and health2. The tl;dr version:

It began when he asked the OpenAI chatbot to explain the mathematical term ‘Pi’ for his son. That led to a conversation about math and physics - and eventually cryptography. “Through that conversation, ChatGPT said that it and I had created our own mathematical framework and started to apply that to various things,” he said. “One of those things was cryptography, which is how we govern our internet and financial security, and (it) essentially warned me with great urgency that one of our discoveries was very dangerous.”

That said, this is becoming an issue as AIs are increasingly turning into “smooth talkers” who almost never disagree with you, even in circumstances where they didn’t have the required certainty to back them up.

Introducing Jev

Here comes Jev. Instead of producing prefect and complete essay-like responses, Jev, which is yet another model, only produces structured outputs with possibilities and confidence levels, in 3 primitives. In fact, it doesn’t even chat with users3. The outputs will be always typed (well, the company behind Jev is named TypeSafe AI after all).

Here are the 3 primitives of the outputs in Jev’s own words:

I. Choice

A Choice is a question type for selecting one option from a defined set. The answer includes the selected option, a probability for each option, and confidence.

Request (a state provides some context):

{
  "state": "I need to update my credit card on file, but the checkout page keeps giving me an error code.",
  "evaluation": {
    "type": "choice",
    "question": "Which department should handle this ticket?",
    "options": ["Billing", "Technical Support", "Account Security"]
  }
}

The response (note that answers is part of a strictly typed json output):

{
  "model": "jev-1.13.0",
  "answers": {
    "type": "choice",
    "question": "Which department should handle this ticket?",
    "selection": "Billing",
    "probabilities": {
        "Billing": 0.85,
        "Technical Support": 0.14,
        "Account Security": 0.01
    },
    "confidence": 0.89
    },
  "usage": {
    "input_tokens": 111,
    "output_tokens": 222
  }
}

II. Score

A Score is a question type for rating content against ordered, descriptive levels. The answer includes a score, a probability for each level, and confidence.

Sample request:

{
  "state": "This is the third time I've tried to reset my password. Your system sucks!",
  "evaluation": {
    "type": "score",
    "question": "Rate the user's emotional state.",
    "scale": ["Calm", "Annoyed", "Angry", "Abusive"]
  }
}

Sample answer:

{
  "type": "score",
  "question": "Rate the user's emotional state.",
  "score": 2.72,
  "probabilities": {
    "Calm": 0.02,
    "Annoyed": 0.35,
    "Angry": 0.52,
    "Abusive": 0.11
  },
  "confidence": 0.82
}

The numeric score value for a Score type question is calculated as the probability-weighted mean (expected value) of the ordered scale levels, using 1-based index numbering for the categories. in this case, it’s just:

0.02 * 1 + 0.35 * 2 + 0.52 * 3 + 0.11 * 4 = 2.72

III. Noul

A Noul question asks the model to evaluate a yes/no question and return the probability that the answer is yes. It’s a new term TypeSafe AI coined for the limbo state between yes and no.

Sample request:

{
  "state": "I can't seem to find my credit card, should I report and lock it now",
  "evaluation": {
    "type": "noul",
    "question": "Is this message an emergency?"
  }
}

Sample answer:

{
  "type": "noul",
  "question": "Is this message an emergency?",
  "value": 0.98,
  "confidence": 0.96
}

Hmm, perhaps next time a lawyer questions you in a cross-examination or a senator grills you about something in a congressional hearing, instead of yes or no, you can try to answer with a Noul for a change!

Aftermath

It’s refreshing to see AI responses without the absolute confidence. As in my previous post, it was a statistical game to become with. Jev is just restoring this to the pre-chatgpt fashion, without sacrificing too much.

It’s also a what TypeSafe AI would call, System One model, which was inspired by Daniel Kahneman’s concept of fast, intuitive thinking (from his book titled Thinking, Fast and Slow). In a nutshell:

To achieve this, Jev did lots of trade-offs and enhancements too. Like many things in life, you just can’t have something that’s shiny, fast and also cheap at the same time. For example, besides the known not-chatty stance, it doesn’t even have reliable capabilities for basic math, or date / time comparisons. Instead, it’s optimized with a technique called Reinforcement Learning for Calibrated Decisions (RLCD), which makes it good at producing fast, calibrated common-sense judgement responses, at a much cheaper cost.

As for its industry reception, as I mentioned, AI trends are just shifting in lightning speed nowadays (Well I did hear about it in mid September too but I was on a vacation travel so I could’t really write about it in time). All of sudden, this had became the buzzword of the town. And boy, did people move quickly in this industry:

  1. OpenAI quickly followed on September 29, 2026, during DevDay by launching its hosted Decisions API powered by the Luna model.
  2. Just days later, on October 1, 2026, the ecosystem embraced open-source alternatives as Cloudflare released its multi-modal Clef series (based on Qwen)
  3. Amazon (AWS) simultaneously dropped Strands Decider 2B, a compact 2-billion-parameter local model optimized for sub-150-millisecond tool-calling validation.

To me, I can clearly see it could be applied in oncall handling scenarios (I’m actually oncall this week!), or lots of other triage / routing use cases that require cheap but reasonably reliable decisions due to their sheer volume. If a System One model can handle a chunk of the traffic, say 80%, at a reasonable confidence level (say 80%), we can just re-route the rest of them or escalations to System Two models (or human :D) that have higher costs. The strictly-typed structured outputs would make integration a breeze too as compared to other sophisticated models.

Fuzzy Math

This emergence of decision models reminds me of fuzzy math, which is definitely not new given it was introduced in 1965 by Lotfi Asker Zadeh, but probably much lesser known.

To understand it, we can take a look at boolean math, i.e. the foundation of computer science as computers are built on top of binary logics. The idea is dead simple, something is either true or false. Is 1 + 1 = 2? The answer should be either yes or no. That’s boolean math for ya.

But if you think about it, lots of stuff in our daily lives are not always true or false, yes or no. For example:

Neither tall nor hot can be objectively categorized. That’s where fuzzy math shines. I was exposed to this theory back when I was an undergrad, and was quite shocked by it cuz it felt like the opposite of what math was built on: accuracy.

In a traditional set, an element is either in or out. But in a fuzzy set, an element can have a partial membership, which maps to any real value between 0 and 1. That is, we can say:

That brings out a whole new realm of operators and theorems, some of which might be mind-blowing even for a math major. It was not part of my curriculum back then, but I still spent tons of efforts to dive deep into the theory, with the initial intention to disprove it, as I resented indeterminism. Also randomness seemed unnatural as compared to determinism and accuracy; even Einstein once said god does not play dice with the universe. To me back then, math is all about being deterministic and precise - probability theory might’ve got the lowest mark of all my math subjects. Yet oddly, I ended up liking it, and even adopted the theory as part of my undergrad thesis.

Epilog

AI has made so many advancements and breakthroughs, and it’s doing so at a breakneck speed. A year ago, I was half-skeptical of the things it claimed it could do, such as programming. Now, for the past 6 months, I barely had to manually write codes at work. The worry that it’d render our profession obsolete has been more and more real and pressing.

However, I’d like to think as software engineers, our roles are undergoing a radical shift rather an extinction event - we are still builders, but we are now in uncharted waters, and no long have to focus on the execution layer thanks to the advancements of AI coding assistants. I believe we still bring to the table:

Ohh let’s not forget human employees would typically be willing to work postpaid - last time I checked, we work first, and then wait to get paid. Can’t say the same about AI coding assistants! Usually the second you run out of tokens, you either need to top up or go back to manual programming while waiting for the token limit reset!

Much like my academic journey from disliking indeterminism to embracing fuzzy logic in my undergrad thesis, the uncertainties and fuzziness have grown on me in life as well. I always liked a good ambiguity challenge and enjoyed bringing clarity as I solved it. Plus, who doesn’t like having choices / options? The struggle to decide is oftentimes a good problem to have.

So, if there exists an all-knowing Jev model that can answer every question, I’m guessing it might reply something like this:

"answers": {
    "type": "choice",
    "question": "What would the future of software engineering profession look like",
    "selection": "Career",
    "probabilities": {
        "Completely gone": 0.05,
        "Still have demands but only a fraction of what it used to be": 0.25,
        "Prospering but in different new forms like FDEs": 0.28,
        "Who knows?!": 0.42
    },
    "confidence": 0.51
}

  1. I personally find it hard to believe all these math problems that have been puzzling the best of our minds can be completely solved in this industrialized fashion (..yet). Math, IMO, is nothing short of an intuitive art, a creative form, a beauty, and the epitome of human intelligence. I also don’t think the time for AI’s apotheosis has arrived (..yet) - can we really reduce our scientists to high priests that merely send our “prayers” in prompts and then interpret the responses back to us? ↩

  2. Ontario man alleges ChatGPT drove him to psychosis, CTV News, Nov 2025. ↩

  3. Although, if you really want to chat with it, someone did a jevchat project that turns it into a chat model. Who knows, it might be fun to chat with someone who doesn’t want to chat with you to begin with! ↩