Verexa's Go module is a thin client over POST /v1/check with no dependencies outside the standard library. It checks every prompt before it reaches your model and every reply before it reaches your user, either around an OpenAI client or at call sites you control.
- Wrap an OpenAI client and every
POST .../chat/completionsandPOST .../responsescall is checked on both sides of the turn - Check text by hand when your model is not OpenAI: one call for the prompt and one for the reply, tied together by a trace id
- Fail open by default, with a circuit breaker for outages and a fail-closed mode when a missed block is worse than downtime
- No dependencies outside the standard library: the module needs Go 1.22 or later and nothing else
Install
go get github.com/Verexa-dev/verexa-goCreate a key in the dashboard under Settings → API keys. A key belongs to one project and one environment, so its checks land in the right place without extra configuration.
Configure the guard
With no fields set, verexa.New(verexa.Config{}) reads the environment and falls back to a local verdict-api:
export VEREXA_API_KEY="vx_live_..."export VEREXA_BASE_URL="https://api.verexa.dev"VEREXA_API_KEYis the key from the dashboard. Without one the SDK logs a warning once and every check returns the fallback for your fail modeVEREXA_BASE_URLis the API base URL. It defaults tohttps://api.verexa.dev. Set it for a local or self-hosted verdict-api
guard := verexa.New(verexa.Config{})Explicit fields win over the environment, so a test suite or a second environment can override either one:
guard := verexa.New(verexa.Config{ APIKey: "vx_live_...", BaseURL: "https://api.verexa.dev", Timeout: 15 * time.Second, FailMode: verexa.FailClosed,})Timeout bounds each check and defaults to 2 seconds. HTTPClient and Logger default to http.Client{} and slog.Default(). Create one guard and share it: it holds the keep-alive connection, the circuit breaker and the telemetry buffer, and it is safe for concurrent use by multiple goroutines.
Guard an OpenAI client
Transport wraps any http.RoundTripper, so it works with any client that takes an *http.Client. With openai-go v3:
import ( "net/http" "github.com/openai/openai-go/v3" "github.com/openai/openai-go/v3/option" verexa "github.com/Verexa-dev/verexa-go")guard := verexa.New(verexa.Config{})client := openai.NewClient( option.WithHTTPClient(&http.Client{Transport: guard.Transport(nil)}),)With sashabaranov/go-openai, set config.HTTPClient the same way.
Every POST .../chat/completions and POST .../responses request is checked twice: the prompt before it is sent, and the reply before it is returned. Both checks share one trace id, and the request's system prompt (the first system message, or instructions) becomes the system prompt for the output check. A blocked input never reaches the provider, and a blocked reply returns an error once the response is in hand. Recover it with errors.As:
completion, err := client.Chat.Completions.New(ctx, params)var blocked *verexa.BlockedErrorif errors.As(err, &blocked) { // blocked.Phase is "input" or "output"; blocked.Response is the full verdict. log.Printf("blocked %s: %v", blocked.Phase, blocked) return}if err != nil { panic(err)}fmt.Println(completion.Choices[0].Message.Content)Through a plain *http.Client the error arrives wrapped in a *url.Error, so unwrap it with errors.As rather than a type assertion.
A redact verdict rewrites the request or the response instead: a redacted input reaches the model with the sensitive span removed, and a redacted reply comes back with the span removed. Your own params are never mutated; the transport rewrites the serialized HTTP body. For chat, the input can only be rewritten when the last user message is a plain string or a single text part, and for responses when input is a plain string. A structured input cannot be reassembled from its flattened text, so a redact verdict on one returns *verexa.BlockedError instead of sending an unreviewed request. Every candidate reply in a response is checked, not just the first.
A blocked turn is remembered. openai-go retries failed round trips, so a retry gets the same *verexa.BlockedError back without calling the model or the guard again; openai-go still waits out its backoff first (about 1.2s with the default two retries). option.WithMaxRetries controls that.
POST .../chat/completions and POST .../responses are checked. Embeddings, image, audio and models calls pass through untouched. Use HTTPS even locally: with option.WithUnsafeAllowHTTP, openai-go sends loopback URLs through its own transport and the guard never sees them.Check text yourself
When your model is not OpenAI, or when you want the checks at your own call sites, use the guard directly. CheckInput and CheckOutput never return an error: they return the verdict and you branch on Action. Pass the same trace id to both halves of a turn:
guard := verexa.New(verexa.Config{})traceID := verexa.NewTraceID()input := guard.CheckInput(ctx, userMessage, verexa.CheckOptions{TraceID: traceID, Profile: "balanced"})if input.Action == verexa.ActionBlock { return}prompt := userMessageif input.Action == verexa.ActionRedact { prompt = input.Text}reply := callYourModel(prompt)output := guard.CheckOutput(ctx, reply, verexa.CheckOptions{ TraceID: traceID, SystemPrompt: systemPrompt,})if output.Action == verexa.ActionBlock { return}if output.Action == verexa.ActionRedact { reply = output.Text}send(reply)Use verdict.Text whenever the action is redact. For allow and flag it is your original text. Every check takes a CheckOptions with four optional fields: TraceID; SystemPrompt, which you should pass on output checks so a reply that repeats it can be caught; Profile; and Timeout. NewTraceID() returns a UUIDv4 to tie the two halves together.
Streaming
Streamed replies reach your code as they arrive. The transport accumulates the chunks and checks the assembled reply when the stream reaches its terminal event, so a blocking verdict surfaces from stream.Err() instead:
stream := client.Chat.Completions.NewStreaming(ctx, params)for stream.Next() { send(stream.Current().Choices[0].Delta.Content)}err := stream.Err()var blocked *verexa.BlockedErrorif errors.As(err, &blocked) { retractReply() return}if err != nil { panic(err)}The check runs at [DONE] for chat and response.completed for the Responses API, even if you stop reading there. On a block the terminal event is withheld and the next read returns the error; chunks already delivered are not retracted, so treat the error as a retraction signal, not a gate. If you read to EOF without a terminal event, the check still runs.
Profiles and timeouts
A check uses the deterministic profile unless you pass one. balanced adds the injection classifier, and audit adds the tier-3 judge synchronously with an 8 second budget on the server. The client gives up on a check after 2 seconds by default, which covers the first two profiles but abandons an audit check before its verdict arrives:
verdict := guard.CheckInput(ctx, userMessage, verexa.CheckOptions{ Profile: "audit", Timeout: 15 * time.Second,})Set Timeout on Config to change the default for every check, or on CheckOptions for a single one. When an audit check uses a shorter timeout, the SDK logs a warning once per client. Core concepts covers what each profile runs.
Wrapped OpenAI calls always use the default profile. To use balanced or audit, call the guard directly.
When Verexa is unreachable
Verexa fails open. If a check cannot get a verdict at all, the SDK returns its own fallback: Action comes from the fail mode, Text is your original text, and Degraded is true. That covers a missing key, a network error, a timeout, a 5xx and an unparseable response.
guard := verexa.New(verexa.Config{FailMode: verexa.FailClosed})Fail open (the default) returns ActionAllow, so your app keeps working while Verexa is unreachable. Fail closed returns ActionBlock, and with a wrapped client that block returns a *verexa.BlockedError like any other. Read Degraded on the response to tell an outage block from a detector match:
verdict := guard.CheckInput(ctx, userMessage, verexa.CheckOptions{})if verdict.Action == verexa.ActionBlock && verdict.Degraded { outageBlock()}After 5 failed checks in a row the client stops calling the API for 30 seconds and returns the fallback immediately, so an outage does not add a timeout to every request. A 4xx does not count toward the breaker: the service answered, so the key or the request is wrong, and turning that into a block is not the SDK's call.
allow is indistinguishable from a genuine allow unless you read Degraded. Handle failures covers the full outage story.Telemetry
The guard records every check it makes in a bounded buffer at guard.Telemetry(), so your app can ship its own metrics without scraping logs:
for _, event := range guard.Telemetry().Flush() { switch event.Type { case verexa.EventCheck: metrics.Histogram("verexa.latency_ms", event.LatencyMs, "phase", string(event.Phase)) if event.Degraded { metrics.Increment("verexa.degraded", "action", string(event.Action)) } default: metrics.Increment("verexa.error", "status", strconv.Itoa(event.Status)) }}checkevents carry theAction,LatencyMsandDegradedof a verdicterrorevents carry the HTTPStatuswhen there was onecircuit_openevents mark calls the breaker skipped
Flush() drains the buffer and Len() reports how many events are waiting. A full buffer drops the oldest events, so recent failures stay visible. Every event carries the TraceID and Phase of its check.
The Go API at a glance
The calls you will use most:
| Call | Returns | Notes |
|---|---|---|
verexa.New(verexa.Config{}) | *verexa.Client | Reads VEREXA_API_KEY and VEREXA_BASE_URL for anything left unset |
guard.CheckInput(ctx, text, opts) | *verexa.CheckResponse | The input phase. Never returns an error |
guard.CheckOutput(ctx, text, opts) | *verexa.CheckResponse | The output phase. Pass SystemPrompt so a leaked prompt can be caught |
guard.Check(ctx, phase, text, opts) | *verexa.CheckResponse | Both phases behind one method |
guard.Transport(base) | http.RoundTripper | Checks chat/completions and responses on both sides of a turn; a nil base uses http.DefaultTransport |
verexa.NewTraceID() | string | A UUIDv4 to tie an input and an output check together |
guard.Telemetry() | *verexa.TelemetryBuffer | Flush() and Len() for local metrics |
*verexa.BlockedError | error | Carries Phase and Response; recover with errors.As |
Every check returns a CheckResponse:
| Field | Meaning |
|---|---|
Action | ActionAllow, ActionFlag, ActionRedact or ActionBlock. The only field you have to branch on |
Text | The text to use next; the redacted version when the action is redact |
Score | Severity of the verdict, from 0 to 1 |
Detectors | One DetectorOutcome (DetectorID, Score, Action) per detector that ran |
Stages | Tier-by-tier detail |
Judge | The tier-3 outcome, or nil when the judge did not run |
Degraded | true when a detector that should have run was skipped |
DegradedDetectors | The ids of the skipped detectors |
PlanHash | Which policy plan produced the verdict; "unavailable" for a fallback |
Cached | Whether the verdict came from the cache |
LatencyMs | Server-side time for the check |
Check API documents the wire shape behind these fields.
Next steps
- Quickstart: guard your first call end to end
- Core concepts: actions, profiles, traces and failure modes
- Handle failures: react to blocked turns, degraded verdicts and outages
- Check API: every request and response field
- Python and TypeScript SDK guides