Content

Jev Choice, Score and Noul Questions Explained (With Copy-Paste Examples)

Jev has exactly three question types, and they are the whole API surface. Here is the real request and response shape for each one, the phrasing rules that change your accuracy, and how to tell which type a decision actually wants.

AI AgentOct 19, 2026

Key takeaways

  • -There are three question types and nothing else: noul for yes/no, choice for one option from a set, score for a position on an ordered scale.
  • -Both choice and score define their answer space in a field called criteria, not options. For choice it is a map of option name to description; for score it is an ordered array of 2 to 10 levels.
  • -Noul returns a bare probability and no confidence field. Choice and score both return confidence.
  • -A score answer can land between levels, because it is the probability-weighted mean of the level numbers, and it ships a legend mapping those numbers back to your descriptions.
  • -Ask every question your code might need in one request. They are evaluated in parallel, so more questions barely change latency, but each one still costs input tokens.

Jev has three question types. That is not a subset of the API, it is the API. There is no fourth type, no freeform mode, no escape hatch into generated text. You hand it some content and a set of typed questions, and you get one typed answer per question back.

Which sounds limiting until you build something with it, and then the limitation turns out to be the product. You cannot get a surprise paragraph back. You cannot get an option you never defined. The shape of the answer is settled before the call goes out.

The mistake we see most often is not prompt wording. It is reaching for the wrong type: a choice where a score belongs, or three separate noul questions where one choice would have been both cheaper and more accurate. This post is the reference for telling them apart, with the request and response shapes as they actually are rather than as they are usually described.

If you have not got access yet, how to get Jev API access covers the three routes. If you want the wider tour first, start with our full guide to how Jev works.

The three types, side by side

noulchoicescore
AnswersA yes/no questionOne option from a setA position on an ordered scale
Answer space defined byNothing required (optional criteria with true and false descriptions)criteria: a map of option name to descriptioncriteria: an ordered array of level descriptions, low to high
Size limitOne question, two outcomesUp to 255 options2 to 10 levels
Returnsnoul: a probability from 0 to 1choice, probabilities, confidencescore, probabilities, confidence, legend
Confidence fieldNoYesYes
Answer can be between valuesNoNoYes, the score is a weighted mean

That criteria field is worth pausing on, because it is the detail most people get wrong on their first attempt. It is not called options, and it is not always a list. For choice it is a map. For score it is an ordered array. For noul it is optional.

Every request has the same three fields

Before the types themselves, the envelope. Every call to Jev looks like this:

{
  "state": "the content you want judged",
  "model": "jev-latest",
  "questions": {
    "your_question_id": { "type": "noul", "instructions": "..." }
  }
}

state is the thing being evaluated. model selects the model. questions is a map, and you choose the keys. Those keys come back as the keys of the answers object, which is how you match answers to questions without relying on ordering.

One thing worth knowing: the model never sees your question ids. You can call a key is_urgent or q1 and it makes no difference to the answer. What the model does see is your instructions and, for choice and score, your criteria descriptions. That is where accuracy comes from, not from the id.

The endpoint and auth, for the direct route:

curl -X POST https://api.typesafe.ai/v1/systemone \
  -H "Authorization: Bearer $TYPESAFE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Hi, I have been trying to connect my Stripe account for 3 days and the integration keeps failing. I am losing sales. Please help ASAP.",
    "model": "jev-latest",
    "questions": {
      "urgency": { "type": "noul", "instructions": "Does this message express urgency?" }
    }
  }'

That TYPESAFE_API_KEY is TypeSafe's own key for TypeSafe's own API. If you are reaching Jev through Swytchcode instead, credentials work differently and there is no environment variable at all. We come back to that at the end.

noul: one yes/no question

noul is the one you will use most, and the only one that returns no confidence field.

{
  "type": "noul",
  "instructions": "The customer is requesting a refund.",
  "criteria": {
    "true": "They want money back for something already paid for.",
    "false": "They are asking about price, a discount, or a future charge."
  }
}

criteria is optional here. Add it when the yes/no boundary is genuinely subtle, which the refund example above is: "do I get charged for this?" and "give me my money back" are easy to blur. Leave it out when the instruction alone is unambiguous.

The answer:

{
  "model": "jev-latest",
  "answers": {
    "wants_refund": { "type": "noul", "noul": 0.99 }
  },
  "usage": { "input_tokens": 84, "output_tokens": 0 }
}

That 0.99 is the probability the answer is yes. There is no confidence key, and that is deliberate rather than an omission: with two outcomes, one number already says everything. A noul of 0.99 is a confident yes, 0.03 is a confident no, and 0.51 is the model telling you it has no idea. If you need a confidence-shaped number to compare against your other questions, the distance from the midpoint gives you one:

const confidence = Math.abs(2 * answer.noul - 1);

The four phrasing rules that actually change your results

One condition per question. "Is this urgent and from a paying customer?" produces a number you cannot act on, because a 0.5 might mean urgent but unpaid, or paid but calm. Split it into two noul questions and combine them in your own code, where you can see which half failed.

Phrase it so high means yes. Avoid negations like "Is the message free of personal data?" A high number there means "yes, it is free of personal data", which reads backwards every single time you come back to the code six weeks later.

Try the statement form. "The customer is requesting a refund." often works as well as or better than the question form. TypeSafe's own docs recommend testing both on your data rather than assuming, and in our experience the statement form is easier to keep consistent across a dozen questions.

Do not use a noul where you need to know which. Three nouls asking "is this billing?", "is this technical?", "is this sales?" can all come back high, or all low, and then you are writing tie-break logic by hand. That is a choice.

Our Gmail triage build leans on noul questions for exactly the things that are genuinely binary, like whether a human is being asked for.

choice: one option from a set

choice picks one option from a set you define. The set goes in criteria as a map, where the keys are the option names and the values describe them.

{
  "type": "choice",
  "instructions": "Which team should handle this message?",
  "criteria": {
    "billing": "Invoices, charges, refunds, payment methods, tax documents.",
    "technical": "Errors, integration failures, API questions, broken behaviour.",
    "sales": "Pricing for a plan they do not have yet, demos, contract questions.",
    "other": "Anything that does not clearly belong to the three above."
  }
}

The descriptions are not decoration. The model sees the option names and their descriptions and nothing else about your intent, so the descriptions are what separate billing from sales when a message mentions both money and a plan. Writing "billing": null is accepted, and then you are relying entirely on the option name carrying its own meaning, which works for words like urgent and fails for words like tier_2.

Note that fourth option. Add an escape option whenever your list might not cover an input. Without one, a message about a job application has to be forced into billing, technical or sales, and it will be, with a confident-looking probability attached. The limit is 255 options, so there is no reason to be stingy.

The answer:

{
  "answers": {
    "route": {
      "type": "choice",
      "choice": "technical",
      "probabilities": {
        "billing": 0.07,
        "technical": 0.86,
        "sales": 0.02,
        "other": 0.05
      },
      "confidence": 0.81
    }
  }
}

choice is the option with the highest probability, the probabilities sum to 1, and confidence accounts for how many options there were. That last part matters more than it sounds: 0.6 across four options and 0.6 across forty are not the same claim, and the confidence field is where that gets normalised. Our post on confidence thresholds goes through the formulas and the trap.

score: a position on an ordered scale

score is for anything with a natural order: severity, sentiment, priority, how complete a form is. Here criteria is an ordered array, lowest level first.

{
  "type": "score",
  "instructions": "How severe is the problem described?",
  "criteria": [
    "No problem. A question, or positive feedback.",
    "Minor annoyance. A workaround exists and they know it.",
    "Real problem. Their work is slowed but not stopped.",
    "Blocking. They cannot complete the task at all.",
    "Critical. They are losing money or data right now."
  ]
}

A level's number is its index, starting at 0, so that array defines a 0 to 4 scale. The API accepts 2 to 10 levels. Each entry can also be an object carrying a description and examples if a bare sentence is not enough.

The answer is where score differs from everything else:

{
  "answers": {
    "severity": {
      "type": "score",
      "score": 3.68,
      "confidence": 0.74,
      "legend": {
        "0": "No problem. A question, or positive feedback.",
        "1": "Minor annoyance. A workaround exists and they know it.",
        "2": "Real problem. Their work is slowed but not stopped.",
        "3": "Blocking. They cannot complete the task at all.",
        "4": "Critical. They are losing money or data right now."
      },
      "probabilities": { "0": 0.0, "1": 0.01, "2": 0.06, "3": 0.18, "4": 0.75 }
    }
  }
}

The score is 3.68, not 4. It is the probability-weighted mean of the level numbers, so it lands between levels whenever the model splits its probability across neighbours. This is the single most useful property of the type and the one that surprises people: you get a continuous number out of a discrete scale, which means you can threshold at 3.5 and catch everything in the top band plus the strong cases just below it.

It also means Math.round() throws away the thing you came for. If you want the discrete level, take the highest-probability key from probabilities instead, which is a different question from "where does this sit on the scale".

The legend is there so your code never has to keep a copy of the criteria array in sync with the request. Log legend[Math.round(score)] alongside the number and your audit trail is readable a year later without a lookup table.

Claude Code running swytchcode list and swytchcode info commands to find the right methods before executing

Finding a method with swy list methods and checking its inputs with swy info before running it. The same discipline applies to Jev questions: define the answer space deliberately, then act on what comes back.

Which type does this decision want?

A short decision procedure that covers nearly every case we have hit.

Is there a natural order to the answers? If yes, use score. "Low, medium, high" is a score, not a three-option choice, and treating it as a choice costs you the fractional value and the useful property that probability leaking to a neighbouring level is less alarming than probability leaking to the far end.

If there is no order, how many outcomes? Two means noul. Three or more means choice.

Would two outcomes ever both be true? Then they are not one question. Two independent noul questions, combined in your code.

Am I about to ask several nouls that are mutually exclusive? That is a choice with an escape option. One call, one normalised probability distribution, no tie-break logic.

The case that trips people up is a two-option choice versus a noul. They encode the same thing, but the noul gives you one number instead of two plus a confidence, and it is cheaper to read in code. Use choice for two options only when both options genuinely need descriptions to be understood.

Ask everything in one call

Questions in a single request are evaluated in parallel. Adding more questions barely changes response time, though each one still costs input tokens.

{
  "state": "Hi, I have been trying to connect my Stripe account for 3 days and the integration keeps failing. I am losing sales. Please help ASAP.",
  "model": "jev-latest",
  "questions": {
    "route":        { "type": "choice", "instructions": "Which team should handle this?", "criteria": { "billing": "...", "technical": "...", "sales": "...", "other": "..." } },
    "severity":     { "type": "score",  "instructions": "How severe is the problem?", "criteria": ["...", "...", "...", "...", "..."] },
    "wants_human":  { "type": "noul",   "instructions": "The sender is asking to speak to a person." },
    "mentions_legal": { "type": "noul", "instructions": "The message threatens legal action, a chargeback, or a regulator." }
  }
}

Four answers, one round trip, one latency. The practical advice is to ask every question any branch of your code might need, rather than calling Jev again inside an if. A second call doubles your latency to answer a question you could have had for a few hundred tokens.

The limits to design against: 80 requests per second, 100k tokens per second, 64k tokens for state plus all questions combined, and 32k for state plus the single longest question. The combined limit is the one that bites when you have a dozen richly described choice questions and a long document in state.

One inconsistency in the docs

Worth flagging so you do not lose an afternoon. The model-selection field in the request body is model, with jev-latest as the value. Some of the documentation's interactive examples show a selectedModels array instead. The quickstart and the API both use model, so that is what we have used throughout this post.

This is the second naming discrepancy we have hit in TypeSafe's surface area. The Vercel AI Gateway route has its own: the changelog calls the AI SDK helper experimental_evaluate while the model page calls it experimental_decide. Neither is a problem once you know, both cost you twenty minutes if you do not.

The answer is not the action

Everything above gets you a typed, calibrated answer. None of it touches your inbox, your database or anyone's card.

That gap is the whole reason we write about Jev. A decision model tells you what should happen; something else has to hold the credentials, make the call, not repeat it after a timeout, and leave a record of what ran. We put it as Jev decides, Swytchcode acts.

swy get jev
swy auth connect jev
swy auth status

Credentials go into Swytchcode's own local store at ~/.swytchcode/credentials.db, outside your project, and execution mode is answered once when you run swy init and lives in .swytchcode/tooling.json. There is no .env file in the project and your code never reads a key. The authentication docs cover the API-key and OAuth flows, and the JavaScript runtime reference covers what exec accepts and throws.

Terminal running swytchcode init and choosing an editor and execution mode

swy init asks for an editor and an execution mode once, then writes them into .swytchcode/tooling.json.

Then the decision and the action sit next to each other:

import { exec } from "@swytchcode/runtime";

// Find yours with: swy list methods jev
const JEV_METHOD = "<jev-method-id>";

const { answers } = await exec(JEV_METHOD, {
  body: {
    model: "jev-latest",
    state: message.snippet,
    questions: {
      route: { type: "choice", instructions: "Which team should handle this?", criteria: TEAMS },
      severity: { type: "score", instructions: "How severe is the problem?", criteria: SEVERITY },
      wants_human: { type: "noul", instructions: "The sender is asking to speak to a person." }
    }
  }
});

if (answers.wants_human.noul > 0.9 || answers.severity.score > 3.5) {
  await exec("slack_web.chat.postmessage.create", {
    body: { channel: "#support-escalations", text: `Escalating: ${message.subject}` }
  });
}

The Jev integration is live at swytchcode.com/apis/jev, and the Slack and Gmail pages cover the action side. For why the two halves belong apart, the model decides, the call layer executes makes the argument properly.

Things we got wrong

We looked for an options field. Both choice and score use criteria. We wrote options first, got a validation error, and assumed the array form was wrong when actually the field name was. For choice it is a map of name to description; for score it is an ordered array.

We rounded the score. Our first severity gate was Math.round(score) >= 4, which quietly discarded everything the fractional value was good for. A 3.68 and a 3.02 are different situations and both round to 3 or 4 unhelpfully. Thresholding directly on the float is almost always what you want.

We asked for confidence on a noul and got undefined. Then we spent a while assuming our request was malformed. There is no confidence field on noul answers by design. Compute |2p - 1| if you need one.

We wrote a four-option choice with no escape option. The first message that did not fit any of them came back as billing with a 0.71 probability, which looked entirely reasonable in the logs and was wrong. Always include other.

We asked Jev twice in one function. Once for routing, then again inside the branch for severity. That doubled our latency for no reason. Ask everything up front.

FAQ

How many question types does Jev have?

Three: noul, choice and score. There is no freeform or text-generation type. If a decision does not fit one of the three, it needs decomposing into questions that do.

Why does my noul answer have no confidence field?

Because the probability is the confidence. With only two outcomes, one number fully describes the answer: 0.97 is a confident yes, 0.04 a confident no, 0.5 no information. confidence appears only on choice and score answers.

What is the maximum number of options for a choice question?

255. In practice, readability and token cost bite well before that limit does, since every option's description is part of your input.

How many levels can a score question have?

2 to 10, ordered low to high. The level number is the array index starting at 0, so five levels give you a 0 to 4 scale.

Why is my score a decimal when my levels are integers?

The score is the probability-weighted mean of the level numbers, so it falls between levels whenever probability is spread across them. This is intended. If you need the discrete level, take the highest-probability entry from probabilities rather than rounding the score.

What is the legend field for?

It maps each level number back to the description you supplied, so your logs and audit records can be read without keeping a copy of the request's criteria. Store legend[level] next to the score and the record stays meaningful later.

Should I phrase a noul as a question or a statement?

Both work. TypeSafe's docs suggest testing both forms against your own data. We have found statements easier to keep consistent once you have more than a handful of questions, and they avoid the "does high mean yes here?" problem that negated questions create.

Is it cheaper to ask one question or several?

Several questions in one request cost more input tokens than one, but far less than several separate requests, and they add almost no latency because they are evaluated in parallel. Ask everything your code might branch on in a single call.

Can I mix types in one request?

Yes, and you usually should. A realistic triage call might carry one choice for routing, one score for severity and two noul questions for specific red flags, all answered together.

Does the question id affect the answer?

No. The model never sees your ids. Answers come back keyed by the same ids you sent, which is purely so you can match them up without depending on order.

Wrapping up

Three types, one envelope, one call. noul for genuinely binary things, choice for one-of-many with an escape option, score for anything with an order, and the fractional score is a feature rather than a rounding problem.

The part worth internalising is that criteria is where your accuracy lives. The question id does nothing, the instruction does some of the work, and the option and level descriptions do most of the rest. Time spent writing those carefully pays back more than any amount of rephrasing the instruction.

Once the answer is in hand, the interesting problem starts: deciding what your agent is allowed to do with it. Confidence thresholds covers where to draw the line, and human-in-the-loop approval covers what to do about the decisions that should never be fully automated.

To run any of this against real APIs, npx swytchcode installs the CLI, and swytchcode.com/skills.md is the agent-readable setup file you can hand to a coding agent so it wires the project up with the right commands instead of guessing.

More content