Jev Confidence Scores: How to Pick Thresholds Before Your Agent Acts
Jev hands back a number with every answer, and most people pick 0.8 because it looks reasonable. Here is what the number actually measures, why choice confidence means different things at three options and thirty, and how to find your own thresholds from logged outcomes.
Key takeaways
- -Noul questions return no confidence field at all. The probability is the uncertainty; use the absolute distance from 0.5 if you need a confidence-shaped number.
- -Choice confidence is normalised against the number of options, so 0.6 with three options and 0.6 with thirty are not the same claim.
- -Score confidence weights by distance, so probability spread over neighbouring levels costs you less than probability on a far-away level.
- -Pick thresholds from the cost of being wrong, not from the number looking high. Labelling an email and refunding a card deserve different bars.
- -The only way to know your real threshold is to log the answer and the outcome, then bucket accuracy by confidence. Everything before that is a guess.
Every Jev answer comes with a number, and that number is the reason Jev is worth building on. A model that says "this email is urgent" is interesting. A model that says "this email is urgent, and I'm 0.91 sure" is something you can write an if statement against.
Which is where it goes wrong. Almost everyone picks 0.8. It looks like a high bar. It's a round number. It feels careful.
It isn't a decision, though, it's a vibe. 0.8 on a two-option question and 0.8 on a thirty-option question are different claims about the world, and neither tells you whether your agent should be allowed to spend money. This post is about what the number actually measures and how to find yours.
What the number is, per question type
The three question types compute confidence differently, and one of them doesn't return it at all.
Noul: there is no confidence field
This surprises people. A noul question returns a probability and that's it:
{ "is_urgent": { "noul": 0.93 } }There's no confidence key because the probability already carries the uncertainty. 0.93 means "probably true". 0.5 means "no idea". 0.07 means "probably false", which is just as informative as 0.93: it's a confident no.
If you want a confidence-shaped number for a noul, TypeSafe's confidence guide suggests distance from the midpoint:
const confidence = Math.abs(2 * p - 1);That maps 0.5 to 0, and both 0.0 and 1.0 to 1. Use it when you need one consistent "how sure is this" value across mixed question types. Don't use it when you care about direction, because it throws away which way the answer went.
The practical consequence: if (answers.is_urgent.noul >= 0.5) is a decision about truth. if (Math.abs(2 * p - 1) >= 0.8) is a decision about certainty. They're different gates and you usually want both.
Choice: normalised against the number of options
Choice returns probabilities across your options, plus a confidence computed from the top one:
confidence = (p_max − 1/n) / (1 − 1/n)n is the number of options. The 1/n is what chance would give you, so the formula is asking how far above random the top answer sits, scaled so that random is 0 and certainty is 1.
With three options it simplifies to (3·p_max − 1) / 2. Which means:
| Options | p_max | Confidence |
|---|---|---|
| 2 | 0.60 | 0.20 |
| 3 | 0.60 | 0.40 |
| 6 | 0.60 | 0.52 |
| 30 | 0.60 | 0.59 |
| 3 | 0.90 | 0.85 |
| 30 | 0.90 | 0.90 |
Same raw probability, confidence from 0.20 to 0.59, because beating chance at 1-in-30 is a stronger signal than beating chance at 1-in-2.
This is the trap. If you tune a threshold of 0.75 on a six-option category question and then reuse it on a two-option one, you've silently made the second question far stricter: 0.75 confidence on two options needs p_max of about 0.875. Thresholds are per question, not per project. The docs say the same thing, and it's the single most common way people get this wrong.
Only the top probability counts, which is worth knowing because it means confidence can't distinguish a clear winner from a near-tie for second. If that distinction matters, read the probabilities yourself:
const probs = Object.values(answers.category.probabilities).sort((a, b) => b - a);
const margin = probs[0] - probs[1]; // how far ahead the winner isTypeSafe suggests that top-to-second ratio as an alternative signal, and for routing questions where two categories are genuinely adjacent, the margin tells you more than the confidence does.
Score: weighted by distance
Score confidence accounts for where the uncertainty sits:
confidence = max(0, 1 − Σᵢ pᵢ·|i − m| / MAD_unif)m is the most likely level, and MAD_unif is the mean absolute deviation you'd get from a uniform spread, which normalises the result.
The behaviour that matters: probability on neighbouring levels barely dents confidence, and probability on distant levels destroys it. A model that's torn between "High" and "Critical" is still telling you something useful. A model split between "Not urgent" and "Critical" is telling you it has no idea, and the formula reflects that even though both are two-way splits.
So for ordered scales, confidence behaves the way you'd want. You can mostly trust it directly, which isn't true of choice.
What the number is not
It isn't a probability that the answer is correct. It's derived from the output distribution. Calibration means those probabilities track reality on the data Jev was trained and evaluated on. Your tickets are not that data.
It isn't comparable across question types. A choice confidence of 0.7 and a score confidence of 0.7 are different computations. Don't average them. Don't put them through one threshold constant.
It isn't a reason. Simon Willison's point about Jev being a black box applies here: you get a float and no explanation. High confidence with the wrong answer looks exactly like high confidence with the right one, which is why thresholds are a risk control and not a correctness guarantee.
It doesn't catch a badly worded question. If your criteria overlap, Jev will confidently pick one of two options that mean nearly the same thing. The confidence is honest about the distribution and silent about your schema being wrong.

TypeSafe reports Jev's probabilities are calibrated via its RLCD training. Calibration on their evaluation data is the starting point for your thresholds, not a substitute for measuring on yours. Source: typesafe.ai
Start from the cost of being wrong
TypeSafe's guidance is three bands, act, check, don't act, with the boundaries set by stakes rather than by any universal number. The only specific figures the docs give are 0.5 as a floor for routing to a human, above 0.9 for high-stakes actions, and acting at 0.5 for something read-only like showing a balance.
The useful reframing: don't ask "is 0.8 high enough". Ask "what does a wrong answer cost, and who finds out".
| Action | Wrong answer costs | Starting bar |
|---|---|---|
| Add a label to an email | Seconds to undo, nobody sees it | 0.5, or no gate at all |
| File a document into a folder | A minute, reversible | 0.6 to 0.7 |
| Post to a Slack channel | Mild noise, visible to the team | 0.7 |
| RSVP tentative to a meeting | Nothing much | 0.7 |
| Decline a meeting | An apologetic phone call | Don't automate it |
| Refund a payment | Real money, hard to claw back | 0.85 and a policy ceiling |
| Send an email as the user | Reputation, irreversible | Approval gate, not a threshold |
Two patterns in that table worth naming.
Reversibility matters more than accuracy. A labelling agent at 85% accuracy is useful because the 15% costs nothing. A refund agent at 97% accuracy is dangerous because the 3% is money. The threshold should track the undo cost, not the hit rate.
Some actions don't get a threshold at all. Sending mail as someone, declining a client's meeting, deleting anything: there's no confidence number that makes those safe, because the failure isn't "slightly wrong", it's "wrong in a way you can't take back". Those get a human, or they don't get automated. We've written about structuring that in human-in-the-loop approval.
The three-band shape in code
const BANDS = {
// per question, because choice confidence depends on option count
category: { act: 0.75, review: 0.45 },
urgency: { act: 0.70, review: 0.40 },
};
function band(name, confidence) {
const b = BANDS[name];
if (confidence >= b.act) return "act";
if (confidence >= b.review) return "review";
return "ignore";
}Three outcomes, not two. The middle band is the one people leave out, and it's where most of the value is: those are the records a human should see, and they're a small enough slice that a human can actually see them.
If your middle band is 40% of your volume, the questions need rewriting, not the thresholds.
Measuring instead of guessing
Everything above is how to pick a starting number. This is how to find the right one.
Log the answer and the outcome. Every decision, with the confidence and what actually turned out to be correct. The second half is the work: it means a human reviewing a sample, or capturing what they did when the agent handed something over.
await log({
record_id: email.id,
question: "category",
answer: answers.category.choice,
confidence: answers.category.confidence,
p_max: Math.max(...Object.values(answers.category.probabilities)),
n_options: Object.keys(answers.category.probabilities).length,
// filled in later, by a person or by what they did next
actual: null,
});Store n_options alongside the confidence. Six months later when someone adds a category, your old thresholds quietly change meaning, and that column is how you notice.
Bucket accuracy by confidence. Once a few hundred rows have outcomes:
SELECT
width_bucket(confidence, 0, 1, 10) / 10.0 AS band,
count(*) AS n,
avg((answer = actual)::int) AS accuracy
FROM decisions
WHERE question = 'category' AND actual IS NOT NULL
GROUP BY band
ORDER BY band;What you want is accuracy climbing with confidence. If it does, your threshold is wherever accuracy crosses what you can live with:
band n accuracy
0.4 88 0.61
0.5 132 0.72
0.6 180 0.81
0.7 240 0.90
0.8 310 0.96
0.9 410 0.99For labelling, 0.7 is fine. For anything that writes to a customer, 0.9. The table makes that a decision instead of an argument.
Watch for the flat curve. If accuracy is 0.78 at every band, confidence isn't separating good answers from bad ones on your data, and no threshold will save you. That's a signal about your question design, usually overlapping criteria or a missing other option, not about the model.
Recheck after any change. New option, reworded criteria, a different jev-latest under you. Pin the model version (jev-1.13.0 rather than jev-latest) if you want thresholds that stay put between runs.
The part a threshold can't do
A threshold is a property of your code. If someone refactors the branch, lowers the constant, or ships a path that forgets to check it, nothing stops the call.
For anything expensive, you want a second gate that isn't in your code at all. In Swytchcode that's a policy, checked before the call leaves, regardless of what your logic decided or how sure the model was:
{
"id": "big-refunds-need-approval",
"target": ["stripe.refund.create3"],
"when": { "field": "amount", "operator": ">", "value": 50000 },
"action": { "type": "REQUIRES_APPROVAL", "message": "Over 500.00 needs a human" },
"approval_timeout": "2h"
}The division is clean: confidence decides what the agent proposes, the policy decides what runs. One is a model output you tuned; the other is a rule that holds even when the tuning is wrong. Our refund approver build shows both layers together, and a prompt is not a policy makes the general case.

Policies and the method allowlist are checked by the runtime on every exec, which is what makes them independent of whatever your threshold decided.
Three more things worth having below the threshold:
- A method allowlist. Only what you added to
tooling.jsoncan run, so a mis-gated decision can't reach a method you never enabled. - Idempotency. A retry after a timeout shouldn't repeat the action. See idempotency explained.
- An audit log.
swy audit networkandswy audit policyare how you reconstruct what a threshold let through at 3am.
Things we got wrong
We used one constant for every question. CONFIDENCE_THRESHOLD = 0.8 across a six-option category question and a two-option noul-style check. The six-option question was gated about right and the two-option one was effectively never allowed to act. Thresholds are per question.
We treated noul's probability as a confidence. if (answers.is_spam.noul > 0.8) reads like a confidence check and isn't: it's "probably spam". A confident not spam (0.03) failed the same check as a genuine coin flip (0.5), so we were treating certainty and truth as the same thing.
We picked 0.8 and never revisited it. For three weeks. When we finally bucketed outcomes, accuracy at 0.65 was already 0.93 for that question, so we'd been sending a third of the easy cases to humans for no reason.
We added a category and kept the threshold. Going from five options to seven changed what 0.75 meant, and the agent got quietly more permissive. Nothing broke loudly, which is the problem.
We averaged confidences across question types. A "combined confidence" from a choice and a score, which is two different formulas added together and means nothing. Gate each answer separately, then combine the decisions.
FAQ
What is a good Jev confidence threshold?
There isn't a universal one. Start from the cost of a wrong answer: around 0.5 for reversible, invisible actions, 0.7 for things a colleague sees, 0.85 and up for anything involving money, and an approval gate rather than a threshold for anything irreversible.
Why does my noul answer have no confidence field?
Because the probability is the uncertainty. Use |2p − 1| if you need a confidence-shaped value.
Why is confidence low when the top probability looks high?
Choice confidence is normalised by option count: (p_max − 1/n) / (1 − 1/n). With two options, p_max of 0.6 is only 0.2 confidence, because 0.5 is chance.
Can I use the same threshold across questions?
No. Choice confidence depends on the number of options, so the same number is a different bar on each question.
Does high confidence mean the answer is right?
No. It means the output distribution was concentrated. Calibration makes that a useful signal on data like Jev's evaluation set; your data is what you have to measure.
How many labelled examples do I need to set a threshold?
A few hundred with known outcomes is enough to see the shape. Thousands if you want tight boundaries in the high band, since that's where the errors are rare.
What if accuracy doesn't rise with confidence?
Your questions are probably the problem: overlapping criteria, no other option, or a question the state can't actually answer. Fix the schema before touching thresholds.
Should I pin the model version?
Yes, if your thresholds matter. jev-latest can move under you; jev-1.13.0 won't.
Wrapping up
The number is the best thing about Jev and the easiest thing to use carelessly. Three facts cover most of it: noul gives you a probability and no confidence, choice confidence is scaled by how many options you offered, and score confidence knows the difference between near-miss and no-idea.
After that it's a cost question, not a model question. Pick a starting bar from what a wrong answer costs, build a three-band gate so unsure cases reach a person, log answers against outcomes, and move the bar when the data tells you to.
And for anything you'd mind getting wrong, put a rule underneath that doesn't depend on the model being right at all.
The Jev guide covers how the question types work, and the Gmail triage and refund approver builds show thresholds doing real work at both ends of the risk scale.
Jev is live as a Swytchcode integration, so swy get jev followed by swy auth connect jev is enough to start running the thresholds in this post against your own account.
More content
OpenAI Dots vs Instinct: Two Ways to Build a Personal Agent
Two personal-agent products launched within a week of each other and solved the same problem differently. Dots gives an agent its own computer inside your workspace; Instinct sends yours into other people's group chats. Both had to invent a permission layer, and neither exposes one to developers.
Jev Pricing Explained: What 1 Million Agent Decisions Actually Cost
A million Jev decisions at 500 input tokens each costs $21.00, and the actions they trigger cost more than the thinking does. Here is the real arithmetic, why free output changes how you write prompts, the break-even against your current model, and the rate limit that actually caps you.
How to Let Jev Triage Your Google Calendar Invites (Accept, Decline, or Ask)
Google Calendar has no accept or decline endpoint. RSVP is an event update that changes your own attendee record, which is a detail most tutorials get wrong. Here is a working invite triage agent with Jev deciding and Swytchcode doing the update.
