Verexa's Python package is a thin, dependency-free client over POST /v1/check. It checks every prompt before it reaches your model and every reply before it reaches your user, either around an OpenAI client or at call sites you control.
- Wrap an OpenAI client and every
chat.completions.createandresponses.createcall is checked on both sides of the turn - Check text by hand when your model is not OpenAI: one call for the prompt and one for the reply, tied together by a trace id
- Fail open by default, with a circuit breaker for outages and a fail-closed mode when a missed block is worse than downtime
- No runtime dependencies: the package needs Python 3.10 or later and nothing else
Install
pip install verexaCreate a key in the dashboard under Settings → API keys. A key belongs to one project and one environment, so its checks land in the right place without extra configuration. The package ships type hints, so editors and mypy --strict see the same API the runtime does.
Configure the guard
With no arguments, create_guard() reads the environment and falls back to a local verdict-api:
export VEREXA_API_KEY="vx_live_..."export VEREXA_BASE_URL="https://api.verexa.dev"VEREXA_API_KEYis the key from the dashboard. Without one the package warns once and every check returns the fallback for your fail modeVEREXA_BASE_URLis the API base URL. It defaults tohttps://api.verexa.dev. Set it for a local or self-hosted verdict-api
from verexa import create_guardguard = create_guard()Explicit arguments win over the environment, so a test suite or a second environment can override either one:
guard = create_guard( api_key="vx_live_...", base_url="https://api.verexa.dev", timeout_ms=15000, fail_mode="closed",)Create one guard and share it. It holds the keep-alive connection, the circuit breaker and the telemetry buffer, and those are safe to use across threads. check_async runs the same client code in a worker thread, so sync and async call sites share all of it.
Guard an OpenAI client
wrap_openai patches chat.completions.create and responses.create in place and returns the same client, so the rest of your code does not change. Every call is checked twice: the prompt before it is sent, and the reply before it is returned. Both checks share one trace id.
from openai import OpenAIfrom verexa import create_guard, wrap_openaiguard = create_guard()client = wrap_openai(OpenAI(), guard)completion = client.chat.completions.create( model="gpt-4o", messages=[{"role": "user", "content": "How do I reset my password?"}],)print(completion.choices[0].message.content)Works the same with AsyncOpenAI:
from openai import AsyncOpenAIfrom verexa import create_guard, wrap_openaiguard = create_guard()client = wrap_openai(AsyncOpenAI(), guard)completion = await client.chat.completions.create( model="gpt-4o", messages=[{"role": "user", "content": "How do I reset my password?"}],)Pass your guard as the second argument when you configured one, as above; without it, wrap_openai builds a client from the environment. guard.wrap_openai(client) does the same. A blocked input raises GuardBlockedError before the provider is called, and a blocked reply raises once the reply is in hand. The error carries the phase and the full verdict:
from verexa import GuardBlockedErrortry: completion = client.chat.completions.create(model="gpt-4o", messages=messages)except GuardBlockedError as err: # err.phase is "input" or "output"; err.response is the full verdict. log.warning("blocked %s: %s", err.phase, [d.detector_id for d in err.response.detectors])A redact verdict rewrites the request or the response instead: a redacted input reaches the model with the sensitive span removed, and a redacted reply comes back with [redacted ...] markers. Your own message list is never mutated. For responses.create, a string input and its instructions are checked, and instructions becomes the system prompt for the output check; a structured input list cannot be reassembled from its flattened text, so a redact verdict on one raises GuardBlockedError instead of sending an unreviewed request.
Check text yourself
When your model is not OpenAI, or when you want the checks at your own call sites, use the guard directly. check_input and check_output never raise on a block: they return the verdict and you branch on action. Pass the same trace id to both halves of a turn.
from verexa import create_guard, random_trace_idguard = create_guard()trace_id = random_trace_id()verdict = guard.check_input(user_message, trace_id=trace_id, profile="balanced")if verdict.action == "block": returnreply = call_your_model(verdict.text if verdict.action == "redact" else user_message)checked = guard.check_output(reply, trace_id=trace_id, system_prompt=SYSTEM_PROMPT)if checked.action == "block": returnsend(checked.text if checked.action == "redact" else reply)Use verdict.text whenever the action is redact. For allow and flag it is your original text. Both methods have async forms that run the blocking call in a worker thread:
verdict = await guard.check_input_async(user_message, trace_id=trace_id)checked = await guard.check_output_async(reply, trace_id=trace_id, system_prompt=SYSTEM_PROMPT)Every check accepts four optional keyword arguments: system_prompt, which you should pass on output checks so a reply that repeats it can be caught; trace_id; profile; and timeout_ms.
Streaming
Streamed replies reach your code as they arrive. The wrapper accumulates the chunks and checks the assembled reply when the stream ends, so a blocking verdict raises GuardBlockedError at that point:
from verexa import GuardBlockedErrorstream = client.chat.completions.create(model="gpt-4o", messages=messages, stream=True)try: with stream: for chunk in stream: send(chunk.choices[0].delta.content or "")except GuardBlockedError: retract_reply()By then the chunks may already be on screen, so treat the error as a retraction signal, not a gate. The wrapped stream still works as a context manager, and with AsyncOpenAI as async for.
stream() and parse() helpers and with_raw_response bypass the guard, so run those turns through check_input and check_output yourself.Profiles and timeouts
A check uses the deterministic profile unless you pass one. balanced adds the injection classifier, and audit adds the tier-3 judge synchronously with an 8 second budget on the server. The client gives up on a check after 2 seconds by default, which covers the first two profiles but abandons an audit check before its verdict arrives:
verdict = guard.check_input(user_message, profile="audit", timeout_ms=15000)Set timeout_ms on a single check, or on the guard to change the default for every check. When an audit check uses a shorter timeout, the package warns once. Core concepts covers what each profile runs.
Wrapped OpenAI calls always use the default profile. To use balanced or audit, call the guard directly.
When Verexa is unreachable
Verexa fails open. If a check cannot get a verdict at all, the package returns its own fallback: action comes from the fail mode, text is your original text, and degraded is true. That covers a missing key, a network error, a timeout, a 5xx and an unparseable response.
from verexa import create_guardguard = create_guard(fail_mode="closed")Fail open (the default) returns allow, so your app keeps working while Verexa is unreachable. Fail closed returns block, and with a wrapped client that block raises GuardBlockedError like any other. Read degraded on the response to tell an outage block from a detector match:
verdict = guard.check_input(user_message)if verdict.action == "block" and verdict.degraded: outage_block()After 5 failed checks in a row the client stops calling the API for 30 seconds and returns the fallback immediately, so an outage does not add a timeout to every request. A 4xx does not count toward the breaker: the service answered, so the key or the request is wrong, and turning that into a block is not the SDK's call.
allow is indistinguishable from a genuine allow unless you read degraded. Handle failures covers the full outage story.Telemetry
The guard records every check it makes in a bounded buffer at guard.telemetry, so your app can ship its own metrics without scraping logs:
for event in guard.telemetry.flush(): if event.type == "check": metrics.histogram("verexa.latency_ms", event.latency_ms, phase=event.phase) if event.degraded: metrics.increment("verexa.degraded", action=event.action) else: metrics.increment("verexa.error", status=event.status or 0)checkevents carry theaction,latency_msanddegradedof a verdicterrorevents carry the HTTPstatuswhen there was onecircuit_openevents mark calls the breaker skipped
flush() drains the buffer and size() reports how many events are waiting. A full buffer drops the oldest events, so recent failures stay visible. Every event carries the trace_id and phase of its check.
The Python API at a glance
Everything the package exports:
| Call | Returns | Notes |
|---|---|---|
create_guard(api_key=None, base_url=None, timeout_ms=None, fail_mode=None) | Guard | Reads VEREXA_API_KEY and VEREXA_BASE_URL for anything left out |
wrap_openai(client, guard=None) | the same client | Checks chat.completions.create and responses.create on both sides of the turn |
guard.check_input(text, ...) | CheckResponse | The input phase. Never raises on a block |
guard.check_output(text, ...) | CheckResponse | The output phase. Pass system_prompt so a leaked prompt can be caught |
guard.check_input_async(...) and check_output_async(...) | awaitable CheckResponse | The same checks, run in a worker thread |
guard.wrap_openai(client) | the same client | Wraps with this guard's configuration |
random_trace_id() | str | A UUID4 to tie an input and an output check together |
guard.telemetry | TelemetryBuffer | flush() and size() for local metrics |
Every check returns a CheckResponse:
| Field | Meaning |
|---|---|
action | allow, flag, redact or block. The only field you have to branch on |
text | The text to use next; the redacted version when the action is redact |
score | Severity of the verdict, from 0 to 1 |
detectors | One DetectorOutcome (detector_id, score, action) per detector that ran |
stages | Tier-by-tier detail |
judge | The tier-3 outcome, or None when the judge did not run |
degraded | True when a detector that should have run was skipped |
degraded_detectors | The ids of the skipped detectors |
plan_hash | Which policy plan produced the verdict; "unavailable" for a fallback |
cached | Whether the verdict came from the cache |
latency_ms | Server-side time for the check |
Check API documents the wire shape behind these fields.
Next steps
- Quickstart: guard your first call end to end
- Core concepts: actions, profiles, traces and failure modes
- Handle failures: react to blocked turns, degraded verdicts and outages
- Check API: every request and response field
- TypeScript and Go SDK guides