Outcome verification for agents that change real systems
Your agent said it worked. Prove it.
A 200 response is not proof, and a timeout is not a failure. Conseqa records what an action was supposed to change, then reads the provider afterward and tells you what is actually true.
The refund had already landed. A blind retry would have paid the customer twice; giving up would have stranded them. This is the ending pnpm demo:protected runs end to end.
Lifecycle · runner eu-prod-runner-02UNKNOWN
✓Intent recorded — refund $2,499.00 · pi_3PjaK2
✓Correlation key issued cq_8f21c04ab9e7
!ECONNRESET after the request was sent
!EXECUTION_AMBIGUOUS · outcome DISCOVERING
·GET /v1/refunds ?payment_intent=pi_3PjaK2
·1 candidate · created before the intent was recorded
Finding nothing is not proof that nothing happened — the write may still be settling, or the read may be scoped too narrowly. Conseqa reports Unknown and leaves the action ambiguous. Whether to retry is your decision, made against a stated gap rather than a guess.
Lifecycle · runner eu-prod-runner-02FAILED
✓Intent recorded — refund $2,499.00 · pi_3PjaK2
✓200 OK — re_1PjaKq8f21c04ab9e7
✓RETURNED_SUCCESS · outcome PENDING · the action did happen
·Outcome re-read at the verification window
✕Refund status is canceled — reversed 4s after creation
✕outcome VIOLATED · action still RETURNED_SUCCESS
Evidence · stripe.refund.create
Refund statuscanceled✕
Amount249900 USD✓
Reversed at14:02:12✕
Correlation keycq_8f21c04ab9e7✓
The only real failure in this set, and the API called it a success. The action and its consequence are tracked separately for exactly this: the call worked, the customer still has no money.
Lifecycle · runner eu-prod-runner-02UNKNOWN
✓Intent recorded — refund $2,499.00 · pi_3PjaK2
!Request timed out after 30s
!EXECUTION_AMBIGUOUS · outcome DISCOVERING
!GET /v1/refunds — 503 from the provider
!Retry 3 of 3 · read path still down
!outcome INCONCLUSIVE · re-queued for the next window
Evidence · stripe.refund.create
Refund recordnot readable—
Matches by keynot readable—
Payment intentpi_3PjaK2✓
Correlation keycq_8f21c04ab9e7✓
Here even the read path is down, so there is nothing to reason from at all. The check is re-queued and the verdict stays Unknown. Absence of evidence is never reported as failure.
Ending 02 is the one pnpm demo:protected runs end to end, with a real runner and a real contract build against a mock provider. The other four illustrate the same contract under different provider behaviour. For a run with no mock in it at all, see the transcript below.
Not an illustration
The same thing, actually run.
Sheet 01 above is drawn. This is a transcript. A real agent writes to a real PostgreSQL database, the reply is destroyed in flight, and the runner recovers the row — and the row is still there afterwards, which is the part you can check yourself.
demo:postgres · Neon · ~15s
$ pnpm demo:postgresConseqa · live PostgreSQL demo (your own Neon database)1. Check the demo table ✓ conseqa_demo.store_credits is reachable
✓ lossy proxy on 127.0.0.1:50769 → ep-divine-snow-….neon.tech:5432
2. Start the private runner ✓ runner on http://127.0.0.1:50770 (loopback only)
✓ contract installed · build e264d7886875
· verifier connects as the read-only role, which holds SELECT and nothing else
3. The agent issues a store credit — and the reply is lost ✓ intent recorded before the statement — durable, on disk
✓ correlation key cq_360ab18bdf18fb058340f112
· INSERT INTO conseqa_demo.store_credits … → ep-divine-snow-….neon.tech
! Connection terminated unexpectedly — the agent has an error and no credit id
! was the customer paid? The agent cannot tell. Retrying might pay them twice.
4. The runner asks the database · SELECT id FROM store_credits WHERE correlation_key = 'cq_360ab…' (read-only)
· runner is searching the database for the row the agent never saw
5. Verdict action CONFIRMED_EXECUTED
outcome SATISFIED
credit 7043bcb7-9626-4416-bedb-ed1004739422 recovered=true
checks 5 of 5 passed
✓ /status issued
✓ /customer_id cus_demo_8042
✓ /amount_minor 2499
✓ /currency USD
✓ /correlation_key cq_360ab18bdf18fb058340f112
The credit is really in your database.
The failure is produced, not simulated. The agent's connection runs through a local TCP proxy that forwards the INSERT to the database and then drops the reply and closes the socket, which is what a network does when it fails at the worst moment. TLS terminates at the database, so the proxy carries bytes it cannot read.
That run is not a recording of one good day — press run and produce a fresh one. Or against your own database, in about a minute:
git clone https://github.com/Devrajsinh-Jhala/conseqa-demo
cd conseqa-demo && npm install
cp .env.example .env # a Postgres connection string
npm run demo
The SDK records the intended outcome to disk and issues a correlation key. In required mode the write does not start until that has succeeded, so there is never an action with no record of what it was for.
The write carries the key
Into a column, or into provider metadata that comes back on every read. This is what makes a lost action findable afterwards by identity rather than by guessing from timestamps and amounts.
After the failure
A runner inside your network reads the system of record and matches on that key. Its database role holds SELECT and nothing else, so verification cannot cause the effect it is checking for.
A verdict with its evidence
Verified, Failed, or Unknown, with the fields that decided it. Here five checks passed and the action moved to CONFIRMED_EXECUTED without anything being retried.
Works withOpenTelemetryStripeRazorpayPostgreSQLAny HTTP APINode.js 22+TypeScript
The gap in every agent trace
A timeout is not a failed action.
When the evidence is genuinely insufficient the verdict is Unknown, and it stays Unknown. Absence of evidence is never converted into failure — that one rule is what makes the other verdicts worth acting on.
Traces record what your process observed. They cannot say whether the business change exists in the provider afterward — so the agent retries and double-refunds, or gives up and strands the customer. Both are guesses, and both are expensive.
Conseqa treats the action and its consequence as two separate facts. An action can be CONFIRMED_EXECUTED while its outcome is VIOLATED — a refund that landed and was reversed four seconds later is not a success, and no status code will tell you so.
Verification is read-only and runs on your credentials, inside your network. Conseqa never issues a refund, never retries an action, and never asks a model whether something probably worked.
One loop, four steps
From intention to verified reality.
Conseqa wraps the action you already take. It does not replace your agent, your provider, or the application code you run today.
01
Record the intent
Before the request leaves your process, the SDK durably records the intended outcome and the pinned contract build.
02
Let the agent act
Your code calls the provider, CRM, or database exactly as it does today. Conseqa does not proxy the call.
03
Read the source of truth
A private runner performs a read-only check against exact IDs, amounts, states, and the correlation key it issued.
04
Return an honest verdict
Verified, Failed, Pending, Overdue, or Unknown — with the evidence attached. Missing evidence is never rendered as failure.
What you install
Three pieces. Secrets stay on your side of the line.
Verification happens inside your environment. The hosted control plane receives redacted lifecycle events, allowed evidence, and traces. Provider secrets stay local. Prompt and tool content capture is off by default; any opt-in content needs your redaction policy.
Your application
TypeScript SDK
Wraps consequential actions, stamps a correlation key, and records intent durably before execution begins.
Your environment
Private runner
A container on your own network. Holds the provider credentials, performs read-only checks, stores durable lifecycle events.
Hosted control plane
Engineering dashboard
Correlates traces to outcomes and evidence. Coverage gaps, grouped issues, alerts, deployment drift, offline regressions.
From zero to a verified action
Wrap one action. Keep everything else.
Create your workspace
Free, and seeded with a Sample environment so you can read a verified, a failed, and an unknown outcome before writing any code.
Run the runner beside your app
One container on your own network. Provider credentials never leave it, and it only ever reads. Runner setup
Protect your first action
A few lines around the call you already make. Discovery, verification, and evidence happen on their own. Full guide
refund.ts
import { Conseqa } from "@conseqa/sdk";
import { decorateStripeRefundRequest, stripeRefundSucceeded } from "@conseqa/connector-stripe";
const conseqa = new Conseqa({
runnerUrl: "http://127.0.0.1:4319",
sidecarToken: process.env.CONSEQA_SIDECAR_TOKEN,
contractRegistry,
});
const refund = await conseqa.protect(stripeRefundSucceeded, {
actionKey: `ticket_${ticketId}_refund`,
protectionMode: "required",
intent: { paymentIntentId, amountMinor, currency: "USD" },
execute: ({ providerCorrelationKey }) => {
// Stamps the key into refund metadata and the Idempotency-Key header.
// Metadata comes back on every read, which is what lets the runner find
// this exact refund again when the response never arrives.
const { body, headers } = decorateStripeRefundRequest(
{ paymentIntentId, amountMinor },
providerCorrelationKey,
);
return stripe.refunds.create(body, { idempotencyKey: headers["Idempotency-Key"] });
},
});
required blocks the call if the intent cannot be durably recorded. The wrapper returns your provider's result unchanged and rethrows the same error object, so removing Conseqa is a one-line revert.
Boundaries
What Conseqa does, and what it deliberately does not.
It is
A reliability layer for agents that modify external systems
An independent, read-only outcome verifier
A trace-to-evidence timeline built for engineers
A way to turn an incident into an offline contract regression test
It is not
An agent builder, support suite, or sales platform
A system that initiates refunds or retries actions on its own
A model guessing whether something probably worked
A claim to observe actions your team has not instrumented
Stop trusting the response. Read the system.
Create a workspace, connect a runner, protect your first consequential action.