Go SDK

Guard OpenAI calls or check any text by hand, from Go 1.22 and up.

Verexa's Go module is a thin client over POST /v1/check with no dependencies outside the standard library. It checks every prompt before it reaches your model and every reply before it reaches your user, either around an OpenAI client or at call sites you control.

  • Wrap an OpenAI client and every POST .../chat/completions and POST .../responses call is checked on both sides of the turn
  • Check text by hand when your model is not OpenAI: one call for the prompt and one for the reply, tied together by a trace id
  • Fail open by default, with a circuit breaker for outages and a fail-closed mode when a missed block is worse than downtime
  • No dependencies outside the standard library: the module needs Go 1.22 or later and nothing else

Install

Install
go get github.com/Verexa-dev/verexa-go

Create a key in the dashboard under Settings → API keys. A key belongs to one project and one environment, so its checks land in the right place without extra configuration.

Configure the guard

With no fields set, verexa.New(verexa.Config{}) reads the environment and falls back to a local verdict-api:

Environment
export VEREXA_API_KEY="vx_live_..."export VEREXA_BASE_URL="https://api.verexa.dev"
  • VEREXA_API_KEY is the key from the dashboard. Without one the SDK logs a warning once and every check returns the fallback for your fail mode
  • VEREXA_BASE_URL is the API base URL. It defaults to https://api.verexa.dev. Set it for a local or self-hosted verdict-api
Create the guard
guard := verexa.New(verexa.Config{})

Explicit fields win over the environment, so a test suite or a second environment can override either one:

Configure explicitly
guard := verexa.New(verexa.Config{	APIKey:   "vx_live_...",	BaseURL:  "https://api.verexa.dev",	Timeout:  15 * time.Second,	FailMode: verexa.FailClosed,})

Timeout bounds each check and defaults to 2 seconds. HTTPClient and Logger default to http.Client{} and slog.Default(). Create one guard and share it: it holds the keep-alive connection, the circuit breaker and the telemetry buffer, and it is safe for concurrent use by multiple goroutines.

Guard an OpenAI client

Transport wraps any http.RoundTripper, so it works with any client that takes an *http.Client. With openai-go v3:

Wrap the OpenAI client
import (	"net/http"	"github.com/openai/openai-go/v3"	"github.com/openai/openai-go/v3/option"	verexa "github.com/Verexa-dev/verexa-go")guard := verexa.New(verexa.Config{})client := openai.NewClient(	option.WithHTTPClient(&http.Client{Transport: guard.Transport(nil)}),)

With sashabaranov/go-openai, set config.HTTPClient the same way.

Every POST .../chat/completions and POST .../responses request is checked twice: the prompt before it is sent, and the reply before it is returned. Both checks share one trace id, and the request's system prompt (the first system message, or instructions) becomes the system prompt for the output check. A blocked input never reaches the provider, and a blocked reply returns an error once the response is in hand. Recover it with errors.As:

Catch a blocked turn
completion, err := client.Chat.Completions.New(ctx, params)var blocked *verexa.BlockedErrorif errors.As(err, &blocked) {	// blocked.Phase is "input" or "output"; blocked.Response is the full verdict.	log.Printf("blocked %s: %v", blocked.Phase, blocked)	return}if err != nil {	panic(err)}fmt.Println(completion.Choices[0].Message.Content)

Through a plain *http.Client the error arrives wrapped in a *url.Error, so unwrap it with errors.As rather than a type assertion.

A redact verdict rewrites the request or the response instead: a redacted input reaches the model with the sensitive span removed, and a redacted reply comes back with the span removed. Your own params are never mutated; the transport rewrites the serialized HTTP body. For chat, the input can only be rewritten when the last user message is a plain string or a single text part, and for responses when input is a plain string. A structured input cannot be reassembled from its flattened text, so a redact verdict on one returns *verexa.BlockedError instead of sending an unreviewed request. Every candidate reply in a response is checked, not just the first.

A blocked turn is remembered. openai-go retries failed round trips, so a retry gets the same *verexa.BlockedError back without calling the model or the guard again; openai-go still waits out its backoff first (about 1.2s with the default two retries). option.WithMaxRetries controls that.

Only two endpoints are guarded
POST .../chat/completions and POST .../responses are checked. Embeddings, image, audio and models calls pass through untouched. Use HTTPS even locally: with option.WithUnsafeAllowHTTP, openai-go sends loopback URLs through its own transport and the guard never sees them.

Check text yourself

When your model is not OpenAI, or when you want the checks at your own call sites, use the guard directly. CheckInput and CheckOutput never return an error: they return the verdict and you branch on Action. Pass the same trace id to both halves of a turn:

Check input and output
guard := verexa.New(verexa.Config{})traceID := verexa.NewTraceID()input := guard.CheckInput(ctx, userMessage, verexa.CheckOptions{TraceID: traceID, Profile: "balanced"})if input.Action == verexa.ActionBlock {	return}prompt := userMessageif input.Action == verexa.ActionRedact {	prompt = input.Text}reply := callYourModel(prompt)output := guard.CheckOutput(ctx, reply, verexa.CheckOptions{	TraceID:      traceID,	SystemPrompt: systemPrompt,})if output.Action == verexa.ActionBlock {	return}if output.Action == verexa.ActionRedact {	reply = output.Text}send(reply)

Use verdict.Text whenever the action is redact. For allow and flag it is your original text. Every check takes a CheckOptions with four optional fields: TraceID; SystemPrompt, which you should pass on output checks so a reply that repeats it can be caught; Profile; and Timeout. NewTraceID() returns a UUIDv4 to tie the two halves together.

Streaming

Streamed replies reach your code as they arrive. The transport accumulates the chunks and checks the assembled reply when the stream reaches its terminal event, so a blocking verdict surfaces from stream.Err() instead:

Check a streamed reply when the stream ends
stream := client.Chat.Completions.NewStreaming(ctx, params)for stream.Next() {	send(stream.Current().Choices[0].Delta.Content)}err := stream.Err()var blocked *verexa.BlockedErrorif errors.As(err, &blocked) {	retractReply()	return}if err != nil {	panic(err)}

The check runs at [DONE] for chat and response.completed for the Responses API, even if you stop reading there. On a block the terminal event is withheld and the next read returns the error; chunks already delivered are not retracted, so treat the error as a retraction signal, not a gate. If you read to EOF without a terminal event, the check still runs.

Profiles and timeouts

A check uses the deterministic profile unless you pass one. balanced adds the injection classifier, and audit adds the tier-3 judge synchronously with an 8 second budget on the server. The client gives up on a check after 2 seconds by default, which covers the first two profiles but abandons an audit check before its verdict arrives:

Check with the audit profile
verdict := guard.CheckInput(ctx, userMessage, verexa.CheckOptions{	Profile: "audit",	Timeout: 15 * time.Second,})

Set Timeout on Config to change the default for every check, or on CheckOptions for a single one. When an audit check uses a shorter timeout, the SDK logs a warning once per client. Core concepts covers what each profile runs.

Wrapped OpenAI calls always use the default profile. To use balanced or audit, call the guard directly.

When Verexa is unreachable

Verexa fails open. If a check cannot get a verdict at all, the SDK returns its own fallback: Action comes from the fail mode, Text is your original text, and Degraded is true. That covers a missing key, a network error, a timeout, a 5xx and an unparseable response.

Fail closed
guard := verexa.New(verexa.Config{FailMode: verexa.FailClosed})

Fail open (the default) returns ActionAllow, so your app keeps working while Verexa is unreachable. Fail closed returns ActionBlock, and with a wrapped client that block returns a *verexa.BlockedError like any other. Read Degraded on the response to tell an outage block from a detector match:

Tell an outage block from a real one
verdict := guard.CheckInput(ctx, userMessage, verexa.CheckOptions{})if verdict.Action == verexa.ActionBlock && verdict.Degraded {	outageBlock()}

After 5 failed checks in a row the client stops calling the API for 30 seconds and returns the fallback immediately, so an outage does not add a timeout to every request. A 4xx does not count toward the breaker: the service answered, so the key or the request is wrong, and turning that into a block is not the SDK's call.

A fallback looks like a verdict
A fallback allow is indistinguishable from a genuine allow unless you read Degraded. Handle failures covers the full outage story.

Telemetry

The guard records every check it makes in a bounded buffer at guard.Telemetry(), so your app can ship its own metrics without scraping logs:

Ship the buffer to your metrics
for _, event := range guard.Telemetry().Flush() {	switch event.Type {	case verexa.EventCheck:		metrics.Histogram("verexa.latency_ms", event.LatencyMs, "phase", string(event.Phase))		if event.Degraded {			metrics.Increment("verexa.degraded", "action", string(event.Action))		}	default:		metrics.Increment("verexa.error", "status", strconv.Itoa(event.Status))	}}
  • check events carry the Action, LatencyMs and Degraded of a verdict
  • error events carry the HTTP Status when there was one
  • circuit_open events mark calls the breaker skipped

Flush() drains the buffer and Len() reports how many events are waiting. A full buffer drops the oldest events, so recent failures stay visible. Every event carries the TraceID and Phase of its check.

The Go API at a glance

The calls you will use most:

CallReturnsNotes
verexa.New(verexa.Config{})*verexa.ClientReads VEREXA_API_KEY and VEREXA_BASE_URL for anything left unset
guard.CheckInput(ctx, text, opts)*verexa.CheckResponseThe input phase. Never returns an error
guard.CheckOutput(ctx, text, opts)*verexa.CheckResponseThe output phase. Pass SystemPrompt so a leaked prompt can be caught
guard.Check(ctx, phase, text, opts)*verexa.CheckResponseBoth phases behind one method
guard.Transport(base)http.RoundTripperChecks chat/completions and responses on both sides of a turn; a nil base uses http.DefaultTransport
verexa.NewTraceID()stringA UUIDv4 to tie an input and an output check together
guard.Telemetry()*verexa.TelemetryBufferFlush() and Len() for local metrics
*verexa.BlockedErrorerrorCarries Phase and Response; recover with errors.As

Every check returns a CheckResponse:

FieldMeaning
ActionActionAllow, ActionFlag, ActionRedact or ActionBlock. The only field you have to branch on
TextThe text to use next; the redacted version when the action is redact
ScoreSeverity of the verdict, from 0 to 1
DetectorsOne DetectorOutcome (DetectorID, Score, Action) per detector that ran
StagesTier-by-tier detail
JudgeThe tier-3 outcome, or nil when the judge did not run
Degradedtrue when a detector that should have run was skipped
DegradedDetectorsThe ids of the skipped detectors
PlanHashWhich policy plan produced the verdict; "unavailable" for a fallback
CachedWhether the verdict came from the cache
LatencyMsServer-side time for the check

Check API documents the wire shape behind these fields.

Next steps