Every Verexa integration, from the OpenAI wrapper to a raw HTTP call, comes down to the same few ideas. Learn them once and the SDKs, the API and the dashboard all read the same way.
- A check sends one piece of text to Verexa, either a prompt or a reply
- A verdict comes back with a single action:
allow,flag,redactorblock - A profile decides which detectors run and how long they may take
- A policy turns single detectors off or into monitor mode for a project and environment
- A trace ties the input check and the output check of one turn together
Checks and phases
A check is one call to POST /v1/check. It carries the text, a phase and a trace id. The phase says which side of your model the text is on:
inputis the prompt, checked before it reaches your modeloutputis the reply, checked before it reaches your user
A guarded turn makes two checks. The OpenAI wrappers make both for you. When you call your model some other way, check the input and the output yourself, as in the Quickstart. For output checks, pass your system prompt as well so Verexa can tell when the reply repeats it.
Verdicts and actions
Each check returns a verdict. The action field is the only one you have to read, and text is the text to use next. It is the redacted version when the action is redact and your original text otherwise.
{ "action": "redact", "score": 1, "detectors": [ { "detectorId": "prompt.instruction_override", "score": 0, "action": "allow" }, { "detectorId": "prompt.unicode_obfuscation", "score": 0, "action": "allow" }, { "detectorId": "text.pii", "score": 1, "action": "redact" } ], "text": "my email is [redacted email]", "latencyMs": 0.05, "degraded": false}| Action | Meaning | What the OpenAI wrapper does |
|---|---|---|
allow | Nothing fired | The text passes through unchanged |
flag | Something looks suspicious but is not certain | The turn continues and the verdict is recorded for review |
redact | Sensitive spans were found and replaced | The turn continues with the redacted text |
block | The text must not get through | The call raises a blocked error |
Every detector that runs reports its own action and a score between 0 and 1. The verdict takes the most severe action among them, in the order allow, flag, redact, block. When two detectors pick the same action, the higher score wins. The detectors array keeps every outcome, so you can always see why a verdict came out the way it did.
Detectors
A detector looks for one kind of problem. Each one runs in a fixed phase and has a default action:
- Rule-based detectors match concrete patterns such as an email address, an API key or your own system prompt. They run in microseconds and give the same answer every time
- The injection classifier (
prompt.injection_classifier) is a model that scores prompts for injection and jailbreak attempts. It runs on input only, and it flags rather than blocks - The judge is a language model that reviews verdicts the cheaper layers are unsure about. It never overrules a
redactorblockfrom a rule-based detector, because those are matches, not guesses
See Guardrails for every detector, its phase and what it catches.
Profiles
A profile picks the detection layers for a check and sets its time budget. Pass it with each check. When you leave it out, Verexa uses deterministic.
| Profile | Budget | Layers | Judge |
|---|---|---|---|
deterministic | 30 ms | Six rule-based detectors | Never |
balanced | 1.5 s | Rule-based detectors and the injection classifier | In the background, for uncertain verdicts |
audit | 8 s | Rule-based detectors and the injection classifier | Before answering, for every verdict that is not already certain |
With balanced, the judge runs after Verexa has already answered. Its verdict does not change the response you got, and it lands in the same trace a moment later. With audit, the response waits for the judge, so the judge's verdict counts.
from verexa import create_guardguard = create_guard(timeout_ms=15000)verdict = guard.check_input(user_message, profile="audit")deterministic and balanced, but an audit check can take up to 8 seconds. Raise the timeout to 15 seconds for audit checks, or the SDK abandons the check and fails open before the verdict arrives.The OpenAI wrappers always use the default profile. To use balanced or audit, call the guard directly.
Policies and monitor mode
A policy adjusts a profile for one project and environment. The API key you send decides which project and environment a check belongs to, so the right policy applies without any extra fields. For each detector in the profile, a policy can:
- Disable it so it does not run at all
- Put it in monitor mode so it still runs and is recorded, but its action is capped at
flag. A monitored detector never blocks or redacts
Monitor mode is the safe way to try a detector on real traffic. Watch what it would have caught in the dashboard, then switch it to enforce. See Policies for how to set one up.
Traces
A trace id groups every check that belongs to one conversation turn. The OpenAI wrappers create one per call and use it for both the input and the output check. When you call the guard yourself, create one with randomTraceId() (random_trace_id() in Python, verexa.NewTraceID() in Go) and pass it to both checks.
In the dashboard, Events shows each trace as one row with its checks, the detectors that fired and the action taken. A background judge verdict from the balanced profile is added to the same trace when it finishes.
Degraded verdicts
A verdict has degraded: true when Verexa could not run everything the profile asked for. That happens when the budget runs out before every detector has run, or when the classifier or the judge cannot be reached. Detectors that were skipped are listed in degradedDetectors.
A degraded verdict is still a real verdict, built from the detectors that did run. Verexa does not cache it, so the next check of the same text tries the full set of layers again.
Fail open and fail closed
If the SDK cannot get a verdict at all, it falls back to one of its own. That covers a missing API key, a network error, a timeout and a 5xx from the API. The fallback has degraded: true and an action that depends on the fail mode:
- Fail open (the default) returns
allow, so your app keeps working while Verexa is unreachable - Fail closed returns
block, so no text gets through unchecked
from verexa import create_guardguard = create_guard(fail_mode="closed")After 5 failed checks in a row, the SDK stops calling the API for 30 seconds and returns the fallback straight away, so an outage does not add a timeout to every request. A 4xx response does not count toward this, because it means your key or request is wrong, not that the API is down.
allow looks like a clean verdict to code that only reads action. If you need to know whether a check really ran, read degraded too. See Handle failures.Next steps
- Guardrails: every detector and what it catches
- Policies: turn detectors on or off and run them in monitor mode
- Handle failures: react to blocked turns, degraded verdicts and outages
- Check API: every request and response field