The TypeScript SDK is a thin client over the check endpoint. Wrap the OpenAI client you already have and both sides of every turn are checked automatically, or call the guard yourself for anything that does not go through OpenAI.
- Wrap once, check twice. Every
chat.completions.createandresponses.createcall is checked before the provider sees the prompt and again before the reply reaches your user - One trace id per turn. The input check and the output check share it, so the dashboard shows a turn as one trace
- Rewrite or stop. A redacted input is rewritten before the request leaves your process; a blocked turn raises
GuardBlockedErrorand never reaches the provider - Fail the way you choose. The SDK fails open by default; one option makes an unreachable API block instead
Install
The package has zero runtime dependencies and ships its own types. It is ESM-only and needs Node 18 or later, which has the global fetch it relies on. openai is not a dependency: install it yourself and give the wrapper the client you already construct.
npm install @verexa/sdkCreate a key in the dashboard under Settings → API keys and set it in the environment:
export VEREXA_API_KEY="vx_live_..."export VEREXA_BASE_URL="https://api.verexa.dev"VEREXA_BASE_URL is optional. Without it the SDK talks to https://api.verexa.dev. Set it to point at a local or self-hosted verdict-api.
Configure the guard
createGuard() reads the environment and returns a guard you can reuse. Options passed here win over the environment:
import { createGuard } from "@verexa/sdk";const guard = createGuard({ apiKey: process.env.VEREXA_API_KEY, baseUrl: "https://api.verexa.dev", failMode: "closed", timeoutMs: 2000,});| Option | Default | Notes |
|---|---|---|
apiKey | VEREXA_API_KEY, then GUARD_API_KEY | Sent as Authorization: Bearer. With no key, checks skip the network and follow the fail mode, after one warning per process |
baseUrl | VEREXA_BASE_URL, then GUARD_API_BASE_URL, then https://api.verexa.dev | Trailing slash is stripped |
timeoutMs | 2000 | Per client; every call can override it. Raise it to 15000 for audit |
failMode | "open" | "open" allows when no verdict arrives; "closed" blocks |
Create one guard per process and reuse it. A guard owns its circuit breaker and telemetry buffer, so a guard created per request resets both.
Wrap the OpenAI client
The wrapper installs itself into the client you pass. chat.completions.create is always wrapped, and responses.create is wrapped when the client has it, so existing call sites do not change.
import OpenAI from "openai";import { createGuard } from "@verexa/sdk";const guard = createGuard();const openai = guard.wrapOpenAI(new OpenAI());const completion = await openai.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "my email is jane.doe@example.com" }],});console.log(completion.choices[0].message.content);- The newest user message is checked on input. Only plain-string content is checked; array content, such as multimodal parts, is skipped. With
responses.create, the check runs on the stringinput, or on the flattened text of a structured one - Every choice is checked on output. Not just
choices[0]; one blocked choice fails the whole call - Redaction rewrites in place. A redacted chat message is replaced before the request is sent, and redacted choice content is replaced before the caller sees it. Structured
responses.createinput that needs redaction raises instead, because the flattened text cannot be reassembled - The system prompt is forwarded (the first
role: "system"message, orinstructions) sooutput.system_prompt_leakcan compare it against the reply
guard.wrapOpenAI(client) and guard.middleware() reuse the guard's client, so one circuit breaker and one telemetry buffer cover everything. The standalone wrapOpenAI(client) builds its own client from the environment.
Check text manually
When a call does not go through OpenAI, or you want the verdict in your own control flow, call the guard directly. checkInput and checkOutput take (text, options) and resolve to a verdict; they never throw, even when no answer arrives. Branch on action yourself, and use text when the action is redact.
import { createGuard, randomTraceId } from "@verexa/sdk";const guard = createGuard();const traceId = randomTraceId();const input = await guard.checkInput(userMessage, { traceId, profile: "balanced" });if (input.action === "block") return;const prompt = input.action === "redact" ? input.text : userMessage;const reply = await callYourModel(prompt);const output = await guard.checkOutput(reply, { traceId, systemPrompt: SYSTEM_PROMPT });if (output.action === "block") return;send(output.action === "redact" ? output.text : reply);Both calls accept the same options: traceId to group the turn, systemPrompt for leak detection on the reply, profile to pick the detection plan, and timeoutMs to override the client default for one call.
Stream a reply
A streamed call is checked on input before the provider is called. The output check runs once the stream ends, after every chunk has already reached you.
import { GuardBlockedError } from "@verexa/sdk";try { const stream = await openai.chat.completions.create({ model: "gpt-4o", messages, stream: true, }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); }} catch (err) { if (!(err instanceof GuardBlockedError)) throw err; if (err.phase === "output") retractReply(); // every chunk is already with the user reply("Sorry, I can't help with that.");}The wrapper returns an async generator rather than the provider's Stream, so consume it with for await; Stream helpers are not available on the result. Tokens pass through exactly as they arrive.
redact verdict on a stream has nothing left to rewrite and is ignored.Use with the Vercel AI SDK
The @verexa/sdk/ai entry point exports a middleware for the AI SDK. It checks and redacts the prompt in transformParams, then checks the generated text, for both generateText and streamText.
import { wrapLanguageModel } from "ai";import { openai } from "@ai-sdk/openai";import { guardMiddleware } from "@verexa/sdk/ai";const model = wrapLanguageModel({ model: openai("gpt-4o"), middleware: guardMiddleware(),});guard.middleware() reuses a configured guard's client; the standalone guardMiddleware() builds one from the environment. Streamed text is checked at the end, the same retraction signal as the OpenAI wrapper.
Read a verdict
Every check returns the same verdict shape. The action is the decision to honor:
allow- nothing fired, the text passes throughflag- the turn continues and the check is recorded for reviewredact- continue with the verdict'stext, not the text you sentblock- stop the turn
| Field | Type | Notes |
|---|---|---|
action | "allow" | "flag" | "redact" | "block" | The decision to honor |
score | number | 0..1; the score of the worst detector |
text | string | The text to use next; rewritten when the action is redact |
detectors | DetectorOutcome[] | Every detector that ran, including ones that returned allow |
latencyMs | number | Total check latency |
degraded | boolean | A detector that should have run was skipped |
degradedDetectors | string[] | The detector ids that were skipped |
planHash | string | The plan that ran; "unavailable" on a fallback |
cached | boolean | The verdict was replayed from the cache |
stages | StageOutcome[] | Per-tier trace: cache, tier 1, tier 2, tier 3 |
judge | JudgeOutcome | undefined | Tier-3 result when the judge ran |
The same object is what checkInput and checkOutput resolve to, and what a wrapped call inspects before it rewrites or raises. Core concepts covers the actions and their fields.
Failures and telemetry
A direct check never rejects. A 500, a timeout, a DNS failure or an unparseable response resolves to a fallback verdict: degraded: true, planHash: "unavailable", and an action set by the fail mode. Only the wrappers throw, and only on a block.
import { GuardBlockedError } from "@verexa/sdk";try { const completion = await openai.chat.completions.create({ model: "gpt-4o", messages }); send(completion.choices[0].message.content);} catch (err) { if (!(err instanceof GuardBlockedError)) throw err; log.warn(`blocked on ${err.phase}: ${err.message}`); if (err.phase === "output") retractReply(); reply("Sorry, I can't help with that.");}GuardBlockedError carries phase and the full response verdict. When you fail closed, an outage produces the same error as a policy block, so read response.degraded to tell them apart.
The default fail mode is open. After 5 failed checks in a row, a circuit breaker returns the fallback immediately for 30 seconds before trying again; a 4xx does not count toward it, because the service answered and the key or request is wrong. For audit traffic, raise the timeout: the tier-3 judge has an 8-second server-side budget and the default 2-second client timeout would abort the check before its verdict arrives.
Wrapped calls generate a trace id per turn and do not return it. The guard's telemetry buffer holds the most recent 500 check, error and circuit-open events, each with its trace id, so it is where to find one:
const events = guard.telemetry.flush();for (const event of events) { console.log(event.traceId, event.phase, event.action, event.degraded);}Handle failures covers what each failure shape means and how to respond to it.
Next steps
- Quickstart: from an API key to a guarded call
- Core concepts: actions, profiles, traces and failure modes
- Handle failures: blocked turns, degraded checks and outages
- API reference: the endpoint behind every check
- The Python and Go SDKs