Content

Jev Pricing Explained: What 1 Million Agent Decisions Actually Cost

A million Jev decisions at 500 input tokens each costs $21.00, and the actions they trigger cost more than the thinking does. Here is the real arithmetic, why free output changes how you write prompts, the break-even against your current model, and the rate limit that actually caps you.

AI AgentOct 20, 2026

Key takeaways

  • -Jev charges $0.042 per million input tokens and nothing for output. A million decisions on 500-token inputs costs $21.00.
  • -Because output is free, the only lever on your bill is how much text you send. Adding more questions to the same call costs tokens; adding a second call costs tokens and latency.
  • -The Vercel AI Gateway route lists $0.04 per million input tokens, marginally under TypeSafe's direct $0.042, and both offer $5 of credit to start.
  • -The limit that binds you changes with prompt size: below about 1,250 input tokens per request it is the 80 requests per second cap, above that it is the 100k tokens per second cap.
  • -TypeSafe's headline 444.6x cheaper claim is self-described as likely on the high end. Our own arithmetic puts realistic savings against a cheap model nearer 3x to 15x.
  • -Decisions and actions are metered separately. Swytchcode counts live commands, which is actions that went out, so a million decisions and a million executed actions are different bills.

A million Jev decisions, on inputs of about 500 tokens each, costs $21.00.

That is the answer to the question in the title, and if it is all you came for you can stop here. The fuller answer is that deciding is the cheap half: once those decisions turn into actions that touch Gmail or Stripe, you are paying for execution too, and for a realistic agent the whole thing lands nearer $170 a month. We get to that arithmetic below, along with the parts of your bill nobody quotes and the one limit that will cap your throughput long before your budget does.

We are writing this because "cheap" is doing a lot of work in most coverage of Jev, including TypeSafe's own. The headline comparison figure is 444.6x cheaper than large language models on decision workflows, which TypeSafe themselves describe as likely on the high end. That is an unusually honest caveat from a vendor, and it is worth taking seriously: our own arithmetic does not get anywhere near 444x against a genuinely cheap model. It gets somewhere between 3x and 15x, which is still a real saving and a much more defensible number to put in a planning document.

The actual price

Two numbers:

  • $0.042 per million input tokens, through TypeSafe directly.
  • $0.00 for output tokens. Not discounted. Free.

The second one is the unusual part and it shapes everything downstream. Jev does not generate text, so there is no output to meter. A response carries a probability, maybe a confidence, and for a score question a legend echoing your own level descriptions back. The usage block in every response reports output_tokens, and in practice you are not billed for them.

So your bill is a function of exactly one thing you control: how many tokens you send.

Cost per decision

At $0.042 per million input tokens:

Input tokens per callCost per decisionDecisions per $1
100$0.0000042238,095
250$0.000010595,238
500$0.000021047,619
1,000$0.000042023,810
2,000$0.000084011,905
4,000$0.00016805,952
8,000$0.00033602,976

The per-decision figures are small enough to be hard to reason about, which is why the right-hand column is there. At a 500-token prompt, one dollar buys you roughly 48,000 decisions.

A million decisions, by prompt size

Scaled up, with the Vercel AI Gateway's listed rate alongside for comparison:

Input tokens per call1M decisions, direct ($0.042/M)1M decisions, gateway ($0.040/M)
100$4.20$4.00
250$10.50$10.00
500$21.00$20.00
1,000$42.00$40.00
2,000$84.00$80.00
4,000$168.00$160.00
8,000$336.00$320.00
16,000$672.00$640.00
32,000$1,344.00$1,280.00

Two things to take from this table.

First, at any volume a startup is likely to hit, this is not a line item anyone will ask you about. Ten thousand decisions a day on 500-token inputs is $0.21 a day, or about $77 a year.

Second, the scaling is entirely linear in prompt size, which means the only real cost-control lever is the size of your state. Halving your input halves your bill, exactly. There is no caching discount to engineer around and no output cost to trade against.

TypeSafe's chart comparing Jev's speed and cost with large language models on decision workflows

TypeSafe's own speed and cost comparison. Worth reading alongside their caveat that the headline multiples are likely on the high end.

The two routes, and the free credit

The price differs slightly depending on how you reach Jev. How to get Jev API access covers the routes in full; the pricing-relevant parts:

TypeSafe directly is $0.042 per million input tokens. Console signups opened on September 20, 2026 with $5 of credit attached, and were paused two days later on September 22, so whether you can sign up today depends on when you read this.

Vercel AI Gateway lists typesafe-ai/jev at $0.04 per million input tokens, with a 32K context window and two providers behind it (TypeSafe AI and DigitalOcean). Paid Vercel teams get $5 of AI Gateway credit per month.

That $5 goes further than it sounds:

Input tokens per callDecisions for $5
250476,190
500238,095
1,000119,048
2,00059,524

A quarter of a million decisions on a free monthly credit is enough that a lot of production workloads will simply never generate an invoice. That is a more interesting fact about Jev's pricing than any multiplier against GPT.

The gateway's 2 cents per million discount is not a reason to pick a route. Pick on whether you want a vendor relationship with TypeSafe, whether you are already billing through Vercel, or whether you want the execution side handled too, which is the Swytchcode route.

Free output changes how you should write prompts

Most cost advice for language models is about keeping the model from rambling. None of that applies here, and the incentives invert in two useful ways.

Ask every question in one call. Questions in a single request are evaluated in parallel, so a request with five questions costs barely more latency than one with a single question. It does cost more input tokens, because each question's instructions and criteria are part of your input. But a second request costs you the whole state again plus a second round trip. If your state is 500 tokens and a question adds 40, then four extra questions in the same call cost 160 tokens while a second call costs 540 tokens and doubles your latency.

Spend tokens on criteria, not on state. Option and level descriptions are where Jev's accuracy comes from, and they are cheap and fixed in size. The variable part of your bill is the content you are judging. Trimming an email body to the first 500 characters before sending it saves real money across a million calls; trimming your carefully written option descriptions saves pennies and costs accuracy. Our guide to the question types goes into what those descriptions are doing.

A worked version of the first point. Say you are triaging support mail with a 600-token body and you need routing, severity, and two red-flag checks:

// One call: 600 tokens of state + ~180 tokens of questions = ~780 tokens
// Cost: $0.0000328 per message
const { answers } = await exec(JEV_METHOD, {
  body: {
    model: "jev-latest",
    state: body,
    questions: { route, severity, wants_human, mentions_legal }
  }
});

Versus the version that asks as it goes:

// Four calls: 4 x (600 + ~45) = ~2,580 tokens, and 4 x the latency
// Cost: $0.0001084 per message, 3.3x more for a worse experience
const route = await ask({ route });
const severity = await ask({ severity });
// ...

Same answers, 3.3 times the cost, four times the latency. This is the single most common way to overspend on Jev.

The break-even against your current model

Comparing against a language model is where published numbers get slippery, because model prices change monthly and vendor comparison charts tend to pick an expensive opponent. So rather than quote prices that will be stale by the time you read this, here is the formula and you can drop today's numbers in.

Jev charges input only:

jev_cost = input_tokens x 0.042 / 1,000,000

A language model doing the same classification charges for input plus the tokens it generates, even if all you wanted was one word:

llm_cost = (input_tokens x P_in + output_tokens x P_out) / 1,000,000

On a 500-token prompt with a 20-token answer, here is what various price points work out to as a multiple of Jev:

Model price (input / output per 1M)Cost per decisionMultiple of Jev
$0.10 / $0.40$0.0000582.8x
$0.25 / $1.00$0.0001456.9x
$0.50 / $2.00$0.00029013.8x
$1.00 / $5.00$0.00060028.6x
$3.00 / $15.00$0.00180085.7x

Look up what you currently pay, find the nearest row, and that is your honest saving. Against a genuinely cheap small model you are looking at something like 3x. Against a mid-tier model, 10x to 30x. The 444x figure requires comparing against a frontier model on a workload with a long generated response, which is not what anyone sensible uses for classification.

Which is a good moment to say the obvious: a 3x saving on a line item of $21 is not a reason to adopt anything. If your classification bill is tens of dollars a month, cost is not your argument for Jev. Latency is (70 to 500ms end to end), and so is the guarantee that you cannot get back an answer outside the set you defined. Save the cost argument for the cases where it is actually load-bearing, which is high-volume, tens of millions of decisions a month.

The limit that caps you first

Price is rarely what stops you. The published limits are:

  • 80 requests per second
  • 100,000 tokens per second
  • 64k tokens for state plus all questions combined
  • 32k tokens for state plus the single longest question

Those first two interact in a way worth knowing, because which one binds you depends on your prompt size:

Input tokens per requestCap from token limitCap from request limitBinding limit
500200 req/s80 req/s80 req/s
1,25080 req/s80 req/sboth, exactly
2,00050 req/s80 req/s50 req/s
4,00025 req/s80 req/s25 req/s

The crossover sits at 1,250 input tokens per request. Below that, you are limited by request count and your tokens-per-second headroom goes unused. Above it, every extra token of state directly reduces how many decisions per second you can make.

That has a practical consequence for anyone processing a backlog. At 500 tokens you can do 80 decisions a second, which is 288,000 an hour and about 6.9 million a day. At 4,000 tokens you can do 25 a second, which is 2.2 million a day, and your cost per decision is eight times higher as well. Trimming state buys you throughput and budget at the same time, and the throughput is usually the one you feel.

What is not in the token price

The $21 covers getting answers. It does not cover doing anything with them, and that is where real systems spend their engineering budget rather than their API budget.

Once a decision is made, something has to hold the credentials for Gmail or Stripe or Slack, make the call, not repeat it when the connection drops, check the action against whatever rules you have, and leave a record of what ran. None of that is in a per-token rate.

Since this post is about what a million decisions costs, it is worth putting our own numbers next to Jev's rather than being coy about it. Swytchcode meters live commands, which is executed actions, not decisions:

PlanPriceLive commands / monthConnected end usersLog retentionApproval workflows
Free$0 forever100107 daysNo
Pro$29/mo or $290/yr10,0001,00030 daysNo
Business$149/mo or $1,490/yr1,000,00010,00090 daysYes
EnterpriseCustomCustomCustomCustomYes

Enterprise is a deployment into your own AWS, Azure or GCP account, with SSO, role-based access control, exportable audit logging and your API keys never leaving your environment. The first two live commands do not need an account at all, and the Free plan does not need a card.

The two meters count different things, and that difference is usually large. A decision is a question you asked Jev. A live command is an action that actually went out to a provider. Most agents decide far more often than they act: a triage agent might run a million Jev decisions in a month and only label, reply or escalate on a fraction of them.

So the honest arithmetic for a million-decision month looks like this. If you act on 12% of your decisions, you need 120,000 live commands, which sits inside Business at $149, plus $21 of Jev tokens. Call it $170 all in. If you act on every single decision, a million live commands is still Business, so the number does not change. If you are deciding a million times and acting a hundred times, you are on Free for the execution side and paying $21 for the thinking.

Worth noting what you get on the execution side that has no token equivalent: allow and deny rules on every plan, approval workflows from Business up, and log retention that determines how far back you can answer "why did the agent do that?" Seven days is fine while you are building. Ninety is what you want the first time someone asks about an action from last month.

swy get jev
swy auth connect jev
swy auth status

Credentials live in Swytchcode's own local store at ~/.swytchcode/credentials.db, outside your project, and execution mode is set once by swy init into .swytchcode/tooling.json. No .env file, no key in your code. The authentication docs cover both flows.

Claude Code running swytchcode list and swytchcode info commands to find the right methods before executing

Looking up a method and checking its inputs before executing it. The decision is cheap; the call it triggers is the part with consequences.

The cost that actually hurts is a duplicated action, not a duplicated decision. Re-running a Jev question costs you two thousandths of a cent. Re-running a refund costs you the refund. That asymmetry is why idempotency belongs in the execution path rather than in your handler, which we went through in idempotency explained and, for the MCP case specifically, in where idempotency keys live now.

The Jev integration is live at swytchcode.com/apis/jev.

How to actually reduce the bill

In order of how much they are worth:

Trim state. Linear savings, and it buys throughput too. Send the email body, not the full MIME. Send the first 500 characters, not the thread.

Batch your questions. One call with five questions instead of five calls. Typically a 3x to 4x saving on real workloads, as above.

Do not re-ask inside a branch. Ask everything any branch might need up front.

Filter before you call. The cheapest decision is the one you never make. If a rule in your own code can route 40% of messages with certainty, run it first and you have cut your Jev volume by 40% for free.

Do not bother with caching. There is no prompt-caching discount to exploit, and identical inputs are rare in classification work anyway.

Things we got wrong

We tried to reproduce the 444.6x figure and could not. Our arithmetic tops out around 86x against an expensive frontier model with a short answer. Getting to 444x needs a long generated response in the comparison, which is not how anyone runs classification. TypeSafe does flag the number as likely on the high end, so this is less a correction than a confirmation that their caveat is the honest part.

We budgeted for output tokens. Our first spreadsheet had an output column, because every other model we have billed for has one. It is zero. The usage block still reports output_tokens, which is what led us to assume they were chargeable.

We assumed 80 requests per second was the ceiling. It is not, above 1,250 tokens per request. We sized a backlog job at 80 req/s with 3,000-token inputs and got throttled at about 33, which is exactly what the token limit predicts. The request cap is the headline number and the token cap is the one that actually applied.

We called Jev once per question. Four calls per message, which read cleanly and cost 3.3 times what it should have. Batching the questions was a two-line change.

FAQ

How much does Jev cost per million tokens?

$0.042 per million input tokens through TypeSafe directly, and $0.04 per million through the Vercel AI Gateway. Output tokens are free.

Is Jev free?

There is no permanently free tier, but there is free credit on both routes. TypeSafe's console offered $5 on signup, and paid Vercel teams get $5 of AI Gateway credit per month, which is about 238,000 decisions at 500 input tokens each. Many small workloads never exceed that.

What does a million Jev decisions cost?

$21.00 at 500 input tokens each. $4.20 if your inputs are 100 tokens, $84.00 if they are 2,000. The scaling is exactly linear in input size.

Why are output tokens free?

Jev does not generate text. It returns a typed answer: a probability, a chosen option, a score, a confidence value. There is no generated content to meter, so there is nothing to charge for.

Is Jev actually cheaper than using GPT or Claude for classification?

Yes, but by less than the headline figures suggest. Against a cheap small model on a 500-token prompt you are looking at roughly 3x; against a mid-tier model, 10x to 30x. Use the formula and table above with whatever you currently pay. At low volumes the saving is a rounding error and latency is the better reason to switch.

Does adding more questions to one request cost more?

Yes, because each question's instructions and criteria count as input tokens. But far less than a second request, which re-sends your whole state and doubles your latency. Always batch.

What is the rate limit?

80 requests per second and 100,000 tokens per second. Which one binds depends on prompt size: the crossover is at 1,250 input tokens per request. There are also structural limits of 64k tokens for state plus all questions and 32k for state plus the longest single question.

Is there a prompt-caching discount?

Not that is published. Since the variable part of a classification prompt is the content being judged, there would be little to cache in any case.

Does the Swytchcode route change the price?

TypeSafe's token pricing is what it is. Swytchcode meters separately, by live command, which is an action that actually went out to a provider rather than a question you asked. Our pricing starts at $0 forever for 100 live commands a month, $29 a month for 10,000, and $149 a month for a million. The first two live commands need no account at all.

So what does a million decisions plus the actions really cost?

$21 of Jev tokens, plus whatever your execution volume lands on. Because most agents decide more often than they act, the two numbers diverge: acting on 12% of a million decisions is 120,000 live commands, which fits the $149 Business plan. Total around $170 for the month. Acting on a hundred of them keeps the execution side on the free plan.

What will actually cost me money in production?

Duplicated actions, not duplicated decisions. A repeated Jev question costs a fraction of a cent. A repeated refund costs the refund. Budget your engineering attention accordingly.

Wrapping up

$0.042 per million input tokens, nothing for output, $21 for a million typical decisions, and a $5 monthly credit that covers a quarter of a million of them. At most volumes this is not a cost you will manage, it is a cost you will forget about.

Which means if you are evaluating Jev, price is the least interesting thing about it. The two reasons that hold up are latency, at 70 to 500ms, and the fact that a typed answer space cannot surprise you. Build the cost argument only when you are genuinely in the tens of millions of decisions, where 3x on a real number starts to matter.

What will cost you is the action on the other side of the decision. Confidence thresholds covers deciding when the agent is sure enough to act, and human-in-the-loop approval covers the actions that should never be fully automatic no matter what the number says.

npx swytchcode installs the CLI, and swytchcode.com/skills.md is the agent-readable setup file you can hand to a coding agent so it configures the project properly instead of guessing.

More content