TypeScript SDK

Wrap the OpenAI client you already have and every turn is checked on both sides.

The TypeScript SDK is a thin client over the check endpoint. Wrap the OpenAI client you already have and both sides of every turn are checked automatically, or call the guard yourself for anything that does not go through OpenAI.

  • Wrap once, check twice. Every chat.completions.create and responses.create call is checked before the provider sees the prompt and again before the reply reaches your user
  • One trace id per turn. The input check and the output check share it, so the dashboard shows a turn as one trace
  • Rewrite or stop. A redacted input is rewritten before the request leaves your process; a blocked turn raises GuardBlockedError and never reaches the provider
  • Fail the way you choose. The SDK fails open by default; one option makes an unreachable API block instead

Install

The package has zero runtime dependencies and ships its own types. It is ESM-only and needs Node 18 or later, which has the global fetch it relies on. openai is not a dependency: install it yourself and give the wrapper the client you already construct.

Install the SDK
npm install @verexa/sdk

Create a key in the dashboard under Settings → API keys and set it in the environment:

Environment
export VEREXA_API_KEY="vx_live_..."export VEREXA_BASE_URL="https://api.verexa.dev"

VEREXA_BASE_URL is optional. Without it the SDK talks to https://api.verexa.dev. Set it to point at a local or self-hosted verdict-api.

Configure the guard

createGuard() reads the environment and returns a guard you can reuse. Options passed here win over the environment:

Configure the guard
import { createGuard } from "@verexa/sdk";const guard = createGuard({  apiKey: process.env.VEREXA_API_KEY,  baseUrl: "https://api.verexa.dev",  failMode: "closed",  timeoutMs: 2000,});
OptionDefaultNotes
apiKeyVEREXA_API_KEY, then GUARD_API_KEYSent as Authorization: Bearer. With no key, checks skip the network and follow the fail mode, after one warning per process
baseUrlVEREXA_BASE_URL, then GUARD_API_BASE_URL, then https://api.verexa.devTrailing slash is stripped
timeoutMs2000Per client; every call can override it. Raise it to 15000 for audit
failMode"open""open" allows when no verdict arrives; "closed" blocks

Create one guard per process and reuse it. A guard owns its circuit breaker and telemetry buffer, so a guard created per request resets both.

Wrap the OpenAI client

The wrapper installs itself into the client you pass. chat.completions.create is always wrapped, and responses.create is wrapped when the client has it, so existing call sites do not change.

Wrap the client and send a turn
import OpenAI from "openai";import { createGuard } from "@verexa/sdk";const guard = createGuard();const openai = guard.wrapOpenAI(new OpenAI());const completion = await openai.chat.completions.create({  model: "gpt-4o",  messages: [{ role: "user", content: "my email is jane.doe@example.com" }],});console.log(completion.choices[0].message.content);
  • The newest user message is checked on input. Only plain-string content is checked; array content, such as multimodal parts, is skipped. With responses.create, the check runs on the string input, or on the flattened text of a structured one
  • Every choice is checked on output. Not just choices[0]; one blocked choice fails the whole call
  • Redaction rewrites in place. A redacted chat message is replaced before the request is sent, and redacted choice content is replaced before the caller sees it. Structured responses.create input that needs redaction raises instead, because the flattened text cannot be reassembled
  • The system prompt is forwarded (the first role: "system" message, or instructions) so output.system_prompt_leak can compare it against the reply

guard.wrapOpenAI(client) and guard.middleware() reuse the guard's client, so one circuit breaker and one telemetry buffer cover everything. The standalone wrapOpenAI(client) builds its own client from the environment.

Wrap a client once
The wrapper mutates the client it is given and returns the same object. Wrapping an already wrapped client installs a second pair of checks, so every turn is checked twice.

Check text manually

When a call does not go through OpenAI, or you want the verdict in your own control flow, call the guard directly. checkInput and checkOutput take (text, options) and resolve to a verdict; they never throw, even when no answer arrives. Branch on action yourself, and use text when the action is redact.

Check input and output
import { createGuard, randomTraceId } from "@verexa/sdk";const guard = createGuard();const traceId = randomTraceId();const input = await guard.checkInput(userMessage, { traceId, profile: "balanced" });if (input.action === "block") return;const prompt = input.action === "redact" ? input.text : userMessage;const reply = await callYourModel(prompt);const output = await guard.checkOutput(reply, { traceId, systemPrompt: SYSTEM_PROMPT });if (output.action === "block") return;send(output.action === "redact" ? output.text : reply);

Both calls accept the same options: traceId to group the turn, systemPrompt for leak detection on the reply, profile to pick the detection plan, and timeoutMs to override the client default for one call.

Stream a reply

A streamed call is checked on input before the provider is called. The output check runs once the stream ends, after every chunk has already reached you.

Consume a guarded stream
import { GuardBlockedError } from "@verexa/sdk";try {  const stream = await openai.chat.completions.create({    model: "gpt-4o",    messages,    stream: true,  });  for await (const chunk of stream) {    process.stdout.write(chunk.choices[0]?.delta?.content ?? "");  }} catch (err) {  if (!(err instanceof GuardBlockedError)) throw err;  if (err.phase === "output") retractReply(); // every chunk is already with the user  reply("Sorry, I can't help with that.");}

The wrapper returns an async generator rather than the provider's Stream, so consume it with for await; Stream helpers are not available on the result. Tokens pass through exactly as they arrive.

A retraction signal, not a gate
The output check runs when the stream ends. A blocking verdict raises as you advance past the last chunk, so every chunk is already with the caller: treat the error as a signal to retract. A redact verdict on a stream has nothing left to rewrite and is ignored.

Use with the Vercel AI SDK

The @verexa/sdk/ai entry point exports a middleware for the AI SDK. It checks and redacts the prompt in transformParams, then checks the generated text, for both generateText and streamText.

Guard a language model
import { wrapLanguageModel } from "ai";import { openai } from "@ai-sdk/openai";import { guardMiddleware } from "@verexa/sdk/ai";const model = wrapLanguageModel({  model: openai("gpt-4o"),  middleware: guardMiddleware(),});

guard.middleware() reuses a configured guard's client; the standalone guardMiddleware() builds one from the environment. Streamed text is checked at the end, the same retraction signal as the OpenAI wrapper.

Read a verdict

Every check returns the same verdict shape. The action is the decision to honor:

  • allow - nothing fired, the text passes through
  • flag - the turn continues and the check is recorded for review
  • redact - continue with the verdict's text, not the text you sent
  • block - stop the turn
FieldTypeNotes
action"allow" | "flag" | "redact" | "block"The decision to honor
scorenumber0..1; the score of the worst detector
textstringThe text to use next; rewritten when the action is redact
detectorsDetectorOutcome[]Every detector that ran, including ones that returned allow
latencyMsnumberTotal check latency
degradedbooleanA detector that should have run was skipped
degradedDetectorsstring[]The detector ids that were skipped
planHashstringThe plan that ran; "unavailable" on a fallback
cachedbooleanThe verdict was replayed from the cache
stagesStageOutcome[]Per-tier trace: cache, tier 1, tier 2, tier 3
judgeJudgeOutcome | undefinedTier-3 result when the judge ran

The same object is what checkInput and checkOutput resolve to, and what a wrapped call inspects before it rewrites or raises. Core concepts covers the actions and their fields.

Failures and telemetry

A direct check never rejects. A 500, a timeout, a DNS failure or an unparseable response resolves to a fallback verdict: degraded: true, planHash: "unavailable", and an action set by the fail mode. Only the wrappers throw, and only on a block.

React to a blocked turn
import { GuardBlockedError } from "@verexa/sdk";try {  const completion = await openai.chat.completions.create({ model: "gpt-4o", messages });  send(completion.choices[0].message.content);} catch (err) {  if (!(err instanceof GuardBlockedError)) throw err;  log.warn(`blocked on ${err.phase}: ${err.message}`);  if (err.phase === "output") retractReply();  reply("Sorry, I can't help with that.");}

GuardBlockedError carries phase and the full response verdict. When you fail closed, an outage produces the same error as a policy block, so read response.degraded to tell them apart.

The default fail mode is open. After 5 failed checks in a row, a circuit breaker returns the fallback immediately for 30 seconds before trying again; a 4xx does not count toward it, because the service answered and the key or request is wrong. For audit traffic, raise the timeout: the tier-3 judge has an 8-second server-side budget and the default 2-second client timeout would abort the check before its verdict arrives.

Wrapped calls generate a trace id per turn and do not return it. The guard's telemetry buffer holds the most recent 500 check, error and circuit-open events, each with its trace id, so it is where to find one:

Read the telemetry buffer
const events = guard.telemetry.flush();for (const event of events) {  console.log(event.traceId, event.phase, event.action, event.degraded);}

Handle failures covers what each failure shape means and how to respond to it.

Next steps