It’s Still About Risk-Taking: Knowing When AI Should Act, Ask, or Abstain

A Follow-Up on Trustworthy Agentic AI — and What Jev Adds to the Discussion

In October 2025, I published:

It’s All About Risk-Taking: Why “Trustworthy” Beats “Deterministic” in the Era of Agentic AI [1]

My main statement was:

Agentic AI won’t be 100% right. It must be trustworthy enough to operate under risk.

Instead of asking whether an AI system is always correct, I suggested asking a different question:

Is the system trustworthy enough for the risk level of this specific use case?

That included topics such as blast radius, human-in-the-loop, permissions, observability, evaluation, rollback, and accountability. [1] Almost one year later, I think there is another side of this argument that deserves attention. What happens if an AI system becomes so careful that it no longer makes a clear decision?

And maybe there is an even more fundamental question:

Should every decision inside an AI system be made by a generative LLM at all?

A new model called Jev, introduced by TypeSafe AI in September 2026 and made available in early access, made me think again about this question. [2]

This is a follow-up to my original risk-taking argument.

Table of contents

  1. AI Can Take Too Much Risk — But Also Too Little
  2. Hallucination and Abstention Are Different Problems
  3. Jev Adds an Interesting Architectural Idea
  4. “Can’t Hallucinate” Does Not Mean “Can’t Be Wrong”
  5. Separate Generation From Decision-Making?
  6. Confidence Is Not the Same as Risk
  7. Trustworthy Does Not Mean Cautious
  8. My Updated View
  9. References and additional resources
  10. Transparency Note

1. AI Can Take Too Much Risk — But Also Too Little

When we discuss AI risk, we usually focus on situations where the system acts although it should not.

For example, an AI system may:

  • Hallucinate information
  • Make an unsupported assumption
  • Select the wrong tool
  • Execute an incorrect action
  • Continue although important context is missing

These are obvious risks.

But there is another extreme.

An AI system can become so cautious that it repeatedly answers:

  • “It depends.”
  • “There is not enough information.”
  • “I cannot determine this.”
  • “Ask a human.”

Sometimes this is exactly the correct behavior.

But if a supposedly autonomous system does this too often, much of the value of automation disappears.

So I see two extremes:

Infographic illustrating the balance of risk in automation, with three categories: 'Always Acts,' 'Calibrated,' and 'Always Defers' indicating levels of automation and decision-making based on risk.
  • Too much risk: The AI acts although uncertainty is too high.
  • Too little risk: The AI avoids a decision although enough evidence is available.

This extends my original idea:

Illustration depicting decision-making actions: 'ACT' with a confident robot using a laptop, 'ASK' showing a robot seeking information, and 'ABSTAIN' with a robot indicating unacceptable uncertainty.

Trustworthiness is not only about preventing wrong actions. It is also about making useful decisions inside an acceptable risk boundary.

2. Hallucination and Abstention Are Different Problems

This distinction matters.

A hallucination occurs when an AI system produces unsupported or incorrect information as if it were valid.

Not making a decision is different.

In machine learning, deliberately rejecting a prediction under uncertainty is commonly discussed as selective classification or the reject option. Research in this area studies the trade-off between risk and coverage: a system can reject uncertain cases to reduce errors, but then it also handles fewer cases autonomously. [6]

This leads to a simple model:

ACT — confidence is sufficient.

ASK — more information could resolve the uncertainty.

ABSTAIN — the remaining uncertainty or the consequences of a wrong decision are unacceptable.

The challenge is not to eliminate abstention. The challenge is to calibrate it. An AI system that never abstains can be dangerous. An AI system that always abstains provides little automation value.

So I would extend my original statement:

A cartoon robot on the left showing 'Too Much Risk' with checkmarks and data, representing hasty actions despite high uncertainty.

Trustworthiness is not the elimination of risk. It is the controlled handling of risk.

3. Jev Adds an Interesting Architectural Idea

On September 15, 2026, TypeSafe AI introduced Jev, which the company describes as its first System One Model. [2]

TypeSafe describes the basic interface as:

“unstructured state in, typed probabilistic decisions out” [2]

Jev is not positioned as another general-purpose chatbot.

TypeSafe currently offers three question types: Choice (choose from predefined options), Score (evaluate against an ordered rubric), and Noul (return a probability for a yes/no question). [3]

The launch also got wider attention. AI Weekly summarized Wall Street Journal reporting about Jev and the copycat models that followed. According to that summary, TypeSafe’s CEO said Jev is used by roughly 25% of Fortune 500 companies. That figure comes from the company and is not independently verified. [5][7]

The company calls its training approach:

Reinforcement Learning for Calibrated Decisions — RLCD. [2]

For me, the interesting part is not simply Jev as a product.

It is the architectural idea.

We currently use generative LLMs for many different responsibilities:

  • Understanding context
  • Reasoning
  • Generating code or text
  • Selecting tools
  • Evaluating results
  • Deciding whether to continue

Modern LLMs are highly capable at these tasks, but fundamentally they still produce token sequences.

There is also an efficiency aspect to this architectural idea. Traditional generative LLMs produce output sequentially, token by token. TypeSafe says Jev instead produces structured probabilistic outputs in parallel and is therefore designed to be substantially faster and less expensive for these bounded decision tasks. [2] 

This was also part of a discussion I had with other AI enthusiasts around this concept: if a workflow only needs a bounded decision such as APPROVE, REJECT, or HUMAN_REVIEW, generating a longer natural-language response may consume computation and output tokens that the surrounding software does not actually need.

I would still treat the magnitude of those savings as workload-dependent. TypeSafe reports very large efficiency and cost improvements in its own evaluations, but the company also notes that some of these results may represent the higher end of real-world gains. [2]

Many control-flow decisions inside software have a different shape.

They are questions such as:

Which tool should I use?

Should the agent continue?

Does this result require human review?

Is this operation acceptable?

Which of these predefined actions should happen next?

Those are bounded decisions.

That makes me wonder whether generation and decision-making should always use the same mechanism.

A flowchart illustrating a robot making decisions based on confidence levels. It shows a process where data leads to three options: ACT, ASK, or ABSTAIN, alongside visual representations of a robot and associated icons for each action.

4. “Can’t Hallucinate” Does Not Mean “Can’t Be Wrong”

TypeSafe makes a strong claim about Jev:

Jev “can’t hallucinate.” [2]

This statement needs context.

TypeSafe itself explains that its reported 0% hallucination value is not an empirical measurement. Schema matching is guaranteed by the architecture, so the model cannot produce output outside the predefined structure. [2]

For example, imagine the only valid options are:

APPROVE
REJECT
HUMAN_REVIEW

Jev cannot suddenly return:

Maybe deploy it somewhere else first.

That output does not exist in the schema.

This is valuable.

But it does not mean the decision is necessarily correct.

For example:

Correct decision:
HUMAN_REVIEW
​
Model decision:
APPROVE

The result is perfectly type-safe.

It is still wrong.

So I would distinguish between:

Schema or output hallucination The system produces something outside the allowed structure.

A diagram explaining the differences between not hallucinating and not being wrong in decision-making processes, featuring terms like output hallucination, allowed outputs (approve, reject, human review), and decision error.

Decision error The system returns a valid but incorrect decision.

That distinction is important for trustworthy AI.

Jev does not eliminate risk.

It changes where the risk is located and how software can handle it.

Diagram illustrating decision-making processes in a model. A robot is shown pondering while analyzing allowed schema options: Approve, Reject, and Human Review. It highlights that a valid output (Approve) may not always be the correct decision, emphasizing the difference between valid output and correct decision.

5. Separate Generation From Decision-Making?

This leads to a possible architecture:

A flowchart illustrating a process that includes 'Context/Request', 'Generative LLM', 'Decision Model', and 'Deterministic Code', leading to actions such as 'Act', 'Ask', and 'Abstain'.

This is my own thought experiment, not an official TypeSafe reference architecture.

But the separation is interesting.

The generative model can handle understanding, reasoning, and generation. For decision-heavy workflows, this separation may also reduce unnecessary generated output and therefore lower latency and token-related cost. Deterministic software can then enforce the actual policy by applying risk-specific thresholds to the returned probabilities or confidence values.

For example:

choice = HUMAN_REVIEW
confidence = 0.84

The surrounding application decides what happens next.

This turns uncertainty into something software can explicitly process.

An illustration depicting a robot contemplating decision-making. On the left, various data sources like documents, databases, and code are shown feeding into the robot, who is analyzing information on a laptop. The robot then faces three options based on a confidence meter: to act, to ask for clarification, or to abstain from making a decision.

6. Confidence Is Not the Same as Risk

This is another important distinction.

For Choice and Score questions, Jev returns a confidence value between 0 and 1. A Noul question returns only a probability. [4]

But confidence is not risk.

Risk only emerges when we combine confidence with the consequence of being wrong.

For me, two parameters matter:

Confidence: How certain is the model about the decision?

Blast radius: What happens if the decision is wrong?

For example:

A diagram illustrating the concept 'Confidence is Not Risk', showing a risk assessment matrix with four quadrants labeled ACT, ASK, ASK, and ABSTAIN based on confidence levels and blast radii.

Renaming a local variable and deleting production data should obviously not use the same acceptance threshold.

TypeSafe’s own documentation describes a related approach: high confidence leads to automatic action, medium confidence to caution or confirmation, and low confidence to routing the case to a human or another system. The documentation also says that thresholds should depend on how risky an action is. [4]

This connects directly to the blast-radius argument from my original post. [1]

And this is what I mean by:

calibrated risk

Risk becomes part of the runtime architecture, not only something written in a governance document.

6.1 The Probability Still Needs Deterministic Decision Logic

One important point is easy to miss: a probability or confidence value does not execute a decision by itself.

The surrounding application still needs explicit deterministic logic that decides what should happen at each confidence level. In its documentation, TypeSafe shows this with conditional code: below 0.5, the case goes to a human. For a high-risk action, even a confidence above 0.9 only leads to “proceed with confirmation”. [4]

Conceptually, this could look like:

if confidence >= 0.90:
ACT or continue with confirmation
elif confidence >= medium_threshold:
ASK or verify
else:
ABSTAIN or route to human review

This deterministic layer is important because the model only provides a probabilistic signal. The application defines the actual risk policy.

The value 0.90 should not be interpreted as a universal rule. TypeSafe explicitly says that the correct thresholds depend on the domain, the consequences of being wrong, and the observed performance of the model for the specific use case. [4]

6.2 What a Decision Model Does Not Give You

A separate decision layer is not free. I see two trade-offs.

No natural-language explanation from the decision model itself. In my original post, I listed decision logs and rationales as part of “trustworthy enough.” [1] TypeSafe describes Jev as giving up string generation and returning typed probabilistic decisions instead of generated text. [2] The explanation behind a decision must therefore come from somewhere else: for example, from the way the question is defined, from surrounding application logic, from logged evidence, or from a generative-model step before or after the decision.

Calibration must be tested. TypeSafe says that higher confidence means higher accuracy. [2] But TypeSafe itself recommends starting with conservative thresholds, testing with your own data, and adjusting over time. [4]

So the trust question does not disappear. It moves from “Do I trust the generated text?” to “Do I trust the confidence value for my use case?”

Flowchart illustrating a decision-making process involving LLM reasoning, context, and actions like asking for more information, acting, abstaining, or escalating issues.

7. Trustworthy Does Not Mean Cautious

This is probably my biggest updated takeaway.

In my original post, I connected trustworthiness mainly with mechanisms such as:

  • Guardrails
  • Observability
  • Evaluation
  • Human-in-the-loop
  • Accountability

I still believe all of these matter.

But I would now add another dimension:

A trustworthy system should know when taking a risk is justified.

Being cautious is not automatically trustworthy.

Imagine an autonomous agent that always says:

“Ask a human.”

Its autonomous error rate might be extremely low.

But it is hardly autonomous.

At the other extreme, an agent that always continues may achieve a high automation rate while creating unacceptable operational risk.

So:

Trustworthiness is not maximum confidence.

Trustworthiness is not minimum risk.

Trustworthiness is calibrated behavior under uncertainty.

8. My Updated View

My original statement was:

Agentic AI won’t be 100% right. It must be trustworthy enough to operate under risk.

I would still use exactly that statement.

But today I would add:

Trustworthiness may also require knowing when to act, when to ask, and when to abstain.

Jev also raises another question for me:

Do we really need a generative LLM for every decision inside an AI agent?

Maybe not.

One possible direction for future Agentic AI architectures is to combine different mechanisms:

Generative models where understanding, reasoning, and generation are needed.

Decision models where bounded probabilistic decisions are useful.

Deterministic software where rules and enforcement must remain explicit.

Humans where accountability, judgment, or unacceptable risk requires them.

I have no evidence that this specific architecture will become the dominant approach.

It is my architectural thought experiment.

But it seems more realistic to me than searching for one perfect AI model that does everything.

In my original post, I closed with:

That’s risk-taking with eyes open.

Today I would extend it:

Trustworthy AI is risk-taking with eyes open — and with a system that knows when the risk is too high.

The objective is not to remove uncertainty.

It is to make uncertainty part of the architecture.

Not agents that always decide.

Not agents that are afraid to decide.

But systems that can distinguish between:

ACT. ASK. ABSTAIN.

Maybe that is what trustworthy enough should increasingly mean.

References and additional resources

[1] Thomas Suedbroecker — It’s All About Risk-Taking: Why “Trustworthy” Beats “Deterministic” in the Era of Agentic AI — October 23, 2025 https://suedbroecker.net/2025/10/23/its-all-about-risk-taking-why-trustworthy-beats-deterministic-in-the-era-of-agentic-ai/

[2] TypeSafe AI — Introducing System One Models & Jev — September 15, 2026 https://typesafe.ai/blog/introducing-system-one-models-and-jev

[3] TypeSafe AI Docs — Introduction (question types Choice, Score, Noul) https://docs.typesafe.ai/introduction

[4] TypeSafe AI Docs — Confidence (confidence ranges and risk-based thresholds) https://docs.typesafe.ai/confidence

[5] The Wall Street Journal — Startup TypeSafe AI’s Jev Model Sparks Copycats, Talk of LLM Alternatives — on or before October 2, 2026 (paywall, not read by the author) — https://www.wsj.com/tech/ai/startup-typesafe-ais-jev-model-sparks-copycats-talk-of-llm-alternatives-e39ff57d

[6] Yonatan Geifman, Ran El-Yaniv — Selective Classification for Deep Neural Networks — 2017 https://arxiv.org/abs/1705.08500

[7] AI Weekly (Alexis Dufresne) — TypeSafe says Jev reaches 25% of Fortune 500, trillion tokens daily — October 2, 2026 — https://aiweekly.co/alerts/typesafe-says-jev-reaches-25-of-fortune-500-trillion-tokens-daily

Transparency Note

This post combines documented information about Jev with my own architectural interpretation.

TypeSafe states that Jev cannot hallucinate because its output is constrained to typed structured decisions. TypeSafe also explicitly notes that its 0% hallucination result follows from guaranteed schema matching rather than an empirical error measurement. I therefore do not interpret the claim as evidence that Jev cannot make incorrect decisions. [2]

The basic idea of routing decisions by confidence is not new. It is described in TypeSafe’s documentation [4] and in research on selective classification [6]. My own interpretation is the ACT — ASK — ABSTAIN naming, the link to blast radius and to my 2025 post, and the architecture that separates a generative LLM, a decision model, deterministic code, and human escalation.

I have not tested Jev myself.

All factual statements about Jev in this post are based on the sources listed above.

As always, this post represents my personal perspective and understanding at the time of writing.


Note: This post reflects my own ideas and experience. AI was used as a writing and thinking aid to structure and check the arguments, not to define them.

#AgenticAI, #TrustworthyAI, #AIRisk, #HumanInTheLoop, #AIEngineering, #SoftwareArchitecture, #AIObservability, #Jev, #TypeSafeAI

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Blog at WordPress.com.

Up ↑