How to Build a Stripe Refund Approver With Jev and a Human in the Loop
Most refund requests are obvious. A few are not, and those are the ones that cost money. This build uses Jev to read the request, a policy to hold anything above a threshold for a human, and idempotency so a retry never refunds twice.
Key takeaways
- -Jev reads the refund request and returns a typed decision with a confidence number, in one call. It never touches Stripe.
- -The approval gate is a policy in policies.json, not a branch in your code, so it still holds even when the model is confident and wrong.
- -Refunds are the textbook case for idempotency. Set it in the execution_policy block of manifest.json before you let anything run in production.
- -Stripe's refund methods come through as numbered variants. stripe.refund.create3 refunds a charge or payment intent; confirm yours with swy info before writing code.
- -Never let the agent refund on its own judgment alone. Confidence gates what it proposes; the policy gates what actually happens.
Refund queues follow a pattern. Seven out of ten requests are obvious: the order never shipped, the customer asked within the window, the amount is small. Two are judgement calls. One is the kind where somebody mentions a chargeback and a lawyer in the same sentence, and you very much want a person reading it.
Support teams handle all ten the same way, by hand, because the cost of automating the seven is the risk of automating the one.
This build separates those. Jev reads each request and proposes a decision with a confidence number. A Swytchcode policy decides what's allowed to happen without a human, independent of what Jev said. And idempotency makes sure a network timeout never turns into two refunds.
That last part is why this post exists. A refund agent that's 95% accurate but occasionally double-refunds is worse than no agent at all.
What you'll build
Refund request (email, ticket, form)
│
▼
Jev: 4 questions in one call
│
▼
{ decision, amount_band, mentions_legal, is_repeat_requester } + confidence
│
├─ deny / escalate ──────────> write a note, no Stripe call
├─ approve, low confidence ──> queue for a human
└─ approve, high confidence ─> swytchcode exec stripe.refund.create3
│
▼
policy check: over ₹50,000?
│
├─ yes ──> held for approval in Slack
└─ no ───> refund runs, idempotency key attachedTwo independent brakes. Jev's confidence decides what the agent proposes. The policy decides what actually runs. The second one matters because the first one is a model.
Why two layers and not just a good prompt
We've written about this at length in a prompt is not a policy, and a refund agent is the clearest case for it.
If your only guard is Jev's confidence, then every refund depends on a model being well calibrated on your data, forever. Calibration is good but it isn't a guarantee, and the failure is silent: a confidently wrong approve looks exactly like a confidently right one.
A policy sits outside the model. It reads the actual request about to go to Stripe, checks the amount, and holds it. It doesn't care how sure anything was. That's the property you want on the call that moves money.
What you need
- Node.js 20 or newer.
- A Jev API key. From console.typesafe.ai, or through Vercel's AI Gateway. Our access guide covers the routes. You'll paste it into Swytchcode once, not into your code.
- The Swytchcode CLI.
npx swytchcode, ornpm install -g swytchcode, thenswy login. - A Stripe account in test mode. Do not point the first version of this at live keys. Stripe's test mode makes real refund objects with no real money.
Step 1: Set up the project
mkdir refund-approver && cd refund-approver
npm init -y && npm pkg set type=module
npm install @swytchcode/runtime
swy init --mode=sandbox
swytchcode init asks which editor you use and whether to run in production or sandbox mode. Start in sandbox.
That writes .swytchcode/tooling.json, which holds the mode. Your code never reads or sets it. There's no .env file in this project either: Swytchcode keeps credentials in its own local store at ~/.swytchcode/credentials.db, outside the project folder.
Step 2: Connect Jev and Stripe, and find the right refund method
This is the step where it's worth slowing down, because Stripe's methods don't come through with the names you'd guess.
swy get jev
swy get stripe
swy discover "create a refund" --library stripeThat returns a family of numbered variants rather than one obvious method:
| Canonical ID | What it does |
|---|---|
| stripe.refund.create3 | Refund a charge or payment intent |
| stripe.refund.create4 | Refund a customer charge or payment intent |
| stripe.refund.create5 | Update a specified refund |
| stripe.refund.create6 | Create a refund for a customer balance |
| stripe.issuing.transactions.refund | Refund a test-mode Issuing transaction |
The numbered suffixes come from how the provider's OpenAPI spec is imported: where several operations have overlapping shapes, each gets its own canonical ID. You'll see the same thing in Gmail, where listing messages is gmail.user.messages.get and fetching one is gmail.user.messages.get1.
So don't guess. Check the contract before you write against it:
swy info stripe.refund.create3
swytchcode info prints a method's inputs and outputs, which is how you confirm you picked the right variant.
For refunding an ordinary payment, stripe.refund.create3 is the one. Add only what you need:
swy add method <jev-method-id> # evaluate questions against a state
swy add method stripe.refund.create3 # refund a charge or payment intent
swy auth connect jev # asks for your Jev API key
swy auth connect stripe # asks for your Stripe secret key
swy doctorTwo methods. That's the whole surface this agent can touch. It cannot create charges, update customers, cancel subscriptions or read your balance, because none of those were added to tooling.json. If a bug tries, Swytchcode refuses. Our post on the config files covers how that allowlist works.
Step 3: Turn on idempotency before anything else
Do this now, not later. A refund is the canonical example of a call you must never send twice, and the duplicate doesn't come from a bug in your logic. It comes from a timeout: your request reaches Stripe, Stripe starts processing, the connection drops, and your code has no idea whether it worked.
Swytchcode sets this per integration in the execution_policy block of the integration's manifest.json:
{
"execution_policy": {
"idempotency": {
"mode": "dynamic"
}
}
}mode: none is the default. mode: dynamic generates a key and reuses it across retries of the same call, so a retried refund is recognised by Stripe as the same refund rather than a new one. How Swytchcode retries failed API calls explains the retry behaviour this pairs with, and idempotency explained covers why the key has to live where the retries happen.
If you take one thing from this post and skip the rest, take this one.
Step 4: Write the questions
The state is the customer's message plus the order facts your system already knows. The questions are deliberately narrow.
// src/decisions.js
export function buildQuestions() {
return {
decision: {
type: "choice",
instructions:
"Given the customer's message and the order facts, what should happen to this refund request?",
criteria: {
approve_full: "Clear-cut case for a full refund under a normal policy",
approve_partial: "Some refund is warranted, but not the full amount",
deny: "The request falls outside any reasonable refund policy",
escalate: "Unusual, sensitive, or outside the patterns above",
},
},
mentions_legal: {
type: "noul",
instructions:
"The customer mentions a chargeback, their bank, a lawyer, legal action, " +
"regulators, or threatens public complaint.",
},
claim_is_specific: {
type: "noul",
instructions:
"The message contains specific, checkable detail about what went wrong " +
"(dates, order numbers, error messages) rather than a general complaint.",
},
frustration: {
type: "score",
instructions: "How frustrated does the customer appear?",
criteria: ["Calm", "Mildly annoyed", "Clearly frustrated", "Angry", "Threatening"],
},
};
}
export function requestToState(req) {
return [
`Customer message: ${req.message}`,
"",
`Order total: ${req.currency} ${req.amount / 100}`,
`Order placed: ${req.placedAt}`,
`Days since order: ${req.daysSinceOrder}`,
`Shipped: ${req.shipped ? "yes" : "no"}`,
`Previous refunds for this customer: ${req.priorRefunds}`,
].join("\n");
}Three decisions in there worth explaining:
escalateis a first-class option, not a fallback. Without it, a weird request gets forced into approve or deny. TypeSafe's Choice docs recommend an explicit "none of the above" for exactly this.mentions_legalis asked separately from the decision. It's not a reason to deny; it's a reason for a human to read it. Keeping it as its own question means you can route on it without it polluting the decision.- The order facts go in the state, not in the instructions. Jev is reading one specific request. Days-since-order and shipped-or-not do more work than any amount of policy prose.
Step 5: Call Jev and Stripe
Both are one exec each, because both are Swytchcode integrations.
// src/jev.js
import { exec } from "@swytchcode/runtime";
// Find yours with: swy list methods jev
const JEV_METHOD = "<jev-method-id>";
export async function decide(state, questions, model = "jev-latest") {
return exec(JEV_METHOD, { body: { model, state, questions } });
}// src/stripe.js
import { exec } from "@swytchcode/runtime";
export async function refund({ paymentIntent, amount, reason = "requested_by_customer" }) {
// amount in the smallest currency unit; omit it for a full refund
return exec("stripe.refund.create3", {
body: {
payment_intent: paymentIntent,
...(amount ? { amount } : {}),
reason,
},
});
}No Stripe SDK, no API key in the code, no retry loop and no idempotency header built by hand. Swytchcode attaches the key you saved with swy auth connect stripe, applies the idempotency mode from manifest.json, retries what's safe to retry, and logs the call. The execution pipeline docs walk through the stages.
Step 6: The rules, which matter more than the model
// src/plan.js
const CONFIDENT = 0.85; // high bar: this one moves money
export function planFor(req, a) {
// 1. Anything legal-adjacent goes to a person, whatever Jev decided.
if (a.mentions_legal.noul >= 0.5) {
return { action: "human", reason: "mentions chargeback or legal action" };
}
// 2. Jev's own escape hatch.
if (a.decision.choice === "escalate") {
return { action: "human", reason: "Jev escalated" };
}
// 3. Deny is a message, never a Stripe call.
if (a.decision.choice === "deny") {
return { action: "deny", reason: "outside refund policy" };
}
// 4. An unsure approval is a queued approval.
if (a.decision.confidence < CONFIDENT) {
return {
action: "human",
reason: `low confidence (${a.decision.confidence.toFixed(2)})`,
};
}
// 5. Vague claims don't get automatic money back.
if (a.claim_is_specific.noul < 0.5) {
return { action: "human", reason: "claim lacks checkable detail" };
}
const partial = a.decision.choice === "approve_partial";
return {
action: "refund",
amount: partial ? Math.round(req.amount / 2) : undefined,
reason: a.decision.choice,
};
}Note what rule 1 does: it ignores confidence entirely. Jev can be completely certain the right answer is a full refund, and if the message mentions a chargeback, a person still reads it. Some rules shouldn't be probabilistic.
The threshold of 0.85 is higher than the 0.5 floor TypeSafe's confidence guide suggests for routing to a human, because the cost of being wrong here is money leaving your account. Our post on picking thresholds goes into how to set these from your own data rather than copying ours.
Step 7: The policy, which doesn't trust any of the above
Everything so far is your code. A policy is the layer underneath it, checked by Swytchcode before the call leaves:
swy policy add{
"id": "big-refunds-need-approval",
"target": ["stripe.refund.create3"],
"when": { "field": "amount", "operator": ">", "value": 50000 },
"action": {
"type": "REQUIRES_APPROVAL",
"message": "Refund over 500.00 proposed by the refund agent"
},
"approval_timeout": "2h"
}Any refund over 50,000 minor units now stops and waits for a person, regardless of what your code decided or how sure Jev was. The human approval docs cover routing those to Slack, and policy rules has the full condition and operator schema.
Worth adding a second one while you're here:
{
"id": "refunds-sandbox-only-for-now",
"target": ["stripe.refund.create3"],
"when": { "field": "amount", "operator": ">", "value": 200000 },
"action": { "type": "DENY", "message": "Refunds over 2000.00 are never automatic" }
}A hard ceiling with no approval path. There are amounts that shouldn't be one Slack click away.
Step 8: Run it
// index.js
import { buildQuestions, requestToState } from "./src/decisions.js";
import { decide } from "./src/jev.js";
import { refund } from "./src/stripe.js";
import { planFor } from "./src/plan.js";
import { pendingRequests } from "./src/queue.js"; // your own source
const apply = process.argv.includes("--apply");
const questions = buildQuestions();
for (const req of await pendingRequests()) {
const { answers } = await decide(requestToState(req), questions);
const plan = planFor(req, answers);
console.log(`[${plan.action.toUpperCase()}] ${req.id} ${req.currency} ${req.amount / 100}`);
console.log(` ${plan.reason}`);
if (!apply || plan.action !== "refund") continue;
try {
const res = await refund({ paymentIntent: req.paymentIntent, amount: plan.amount });
console.log(` refunded: ${res.id}`);
} catch (err) {
// A policy hold arrives here too: the call didn't run, and that's correct
console.log(` not refunded: ${err.message}`);
}
}Dry run first:
node index.js[REFUND] req_8821 INR 1499
approve_full
[HUMAN ] req_8822 INR 24990
mentions chargeback or legal action
[REFUND] req_8823 INR 899
approve_partial
[DENY ] req_8824 INR 4500
outside refund policy
[HUMAN ] req_8825 INR 67500
low confidence (0.71)Read a few hundred of these against what your support team actually decided before you pass --apply. The output is the cheap part; agreement with your humans is the thing you're testing.
Make it safer still
Test mode until you're bored of it. Stripe test keys make real refund objects and move no money. There is no reason to be in live mode while you're still tuning question wording.
Check the audit log. swy audit network lists every Stripe call the agent made and swy audit policy shows what was held or denied. After an overnight run, that log answers "what did it refund?" without you having to trust the console output.
Watch the approval queue length. If most refunds end up waiting for a human, the agent isn't saving anyone time. Either the threshold is too high or the questions need better wording. Measure it before you declare success.
Keep escalate honest. If Jev escalates 30% of requests, your criteria probably overlap. Rewrite them rather than lowering the bar.
Things we got wrong the first time
We used stripe.refund.create. It doesn't exist. We assumed the canonical ID would mirror Stripe's REST path and wrote code against a method that was never in the manifest. swy discover and swy info take ten seconds and would have saved the debugging.
We put the amount check in code instead of a policy. It worked, and it was the wrong place. A number in a JavaScript file is a number someone can refactor past. Moving the ceiling into policies.json meant the limit survived changes to the agent, and showed up in swy audit policy when it fired.
We let frustration influence the decision. Our first version nudged toward approving when frustration was high. It took one afternoon to realise we'd built a system that rewards shouting. The score is still in the payload, but it only routes to a human faster. It never changes the answer.
We tested on approved refunds only. Our first sample came from the approved pile, so the agent looked excellent. The requests that teach you something are the denied ones and the weird ones.
Where to go from here
- Post the proposal to Slack instead of refunding. For the first month, have the agent recommend and a human click. You get the measurement without the exposure.
- Feed the decision back. Log Jev's answer and what the human actually did, then compare them by confidence band. That's how you find your real thresholds.
- Add partial-refund reasoning. Right now
approve_partialhalves the amount, which is crude. Ascorequestion over proposed bands is a better fit. - Handle the other direction. A
noulfor "this customer appears to be refund-farming" plus a count of prior refunds catches a pattern no single request shows.
The Jev guide covers eight more use cases, and the examples repo has a working refund agent you can borrow from.
FAQ
Can Jev issue a refund by itself?
No. Jev only returns decisions. Every Stripe call goes through Swytchcode, and only methods you added to tooling.json can run at all.
Which Stripe method refunds a payment?
stripe.refund.create3 refunds a charge or payment intent. Stripe's methods arrive as numbered variants, so confirm with swy discover "create a refund" --library stripe and swy info for your fetched version.
How do I stop a retry from refunding twice?
Set idempotency to mode: dynamic in the execution_policy block of the Stripe integration's manifest.json. The default is none.
What happens when a policy holds a refund?
The call doesn't run. Your exec throws, the request waits for approval for the configured timeout, and the hold is recorded in swy audit policy.
Where does my Stripe secret key go?
Into Swytchcode's credential store via swy auth connect stripe, not your code and not a .env file.
What confidence threshold should I use?
Higher than you'd use for labelling an email. We start at 0.85 for approvals and tune from logged outcomes. See picking thresholds.
Does this work for partial refunds?
Yes. Pass amount in the smallest currency unit; omit it for a full refund.
Can I use Python?
Yes. The Python runtime has the same model, so the same Jev and Stripe calls work from Python.
Wrapping up
The agent is small. The interesting parts are the three places it refuses to act: a legal mention, a low confidence score, and a policy ceiling it can't see past.
Jev is good at reading a refund request, which is genuinely useful and was hard a year ago. It is not the thing that makes automating refunds safe. The allowlist, the approval policy and the idempotency key are, and all three sit outside the model on purpose.
Start with the Stripe integration and the Jev integration, or hand your coding agent our skills file and ask it to build a refund approver with Jev.
More content
OpenAI Dots vs Instinct: Two Ways to Build a Personal Agent
Two personal-agent products launched within a week of each other and solved the same problem differently. Dots gives an agent its own computer inside your workspace; Instinct sends yours into other people's group chats. Both had to invent a permission layer, and neither exposes one to developers.
Jev Confidence Scores: How to Pick Thresholds Before Your Agent Acts
Jev hands back a number with every answer, and most people pick 0.8 because it looks reasonable. Here is what the number actually measures, why choice confidence means different things at three options and thirty, and how to find your own thresholds from logged outcomes.
How to Let Jev Triage Your Google Calendar Invites (Accept, Decline, or Ask)
Google Calendar has no accept or decline endpoint. RSVP is an event update that changes your own attendee record, which is a detail most tutorials get wrong. Here is a working invite triage agent with Jev deciding and Swytchcode doing the update.
