Core concepts

How Verexa decides what happens to a prompt or a reply.

Every Verexa integration, from the OpenAI wrapper to a raw HTTP call, comes down to the same few ideas. Learn them once and the SDKs, the API and the dashboard all read the same way.

  • A check sends one piece of text to Verexa, either a prompt or a reply
  • A verdict comes back with a single action: allow, flag, redact or block
  • A profile decides which detectors run and how long they may take
  • A policy turns single detectors off or into monitor mode for a project and environment
  • A trace ties the input check and the output check of one turn together

Checks and phases

A check is one call to POST /v1/check. It carries the text, a phase and a trace id. The phase says which side of your model the text is on:

  • input is the prompt, checked before it reaches your model
  • output is the reply, checked before it reaches your user

A guarded turn makes two checks. The OpenAI wrappers make both for you. When you call your model some other way, check the input and the output yourself, as in the Quickstart. For output checks, pass your system prompt as well so Verexa can tell when the reply repeats it.

Verdicts and actions

Each check returns a verdict. The action field is the only one you have to read, and text is the text to use next. It is the redacted version when the action is redact and your original text otherwise.

Verdict for a prompt with an email address (trimmed)
{  "action": "redact",  "score": 1,  "detectors": [    { "detectorId": "prompt.instruction_override", "score": 0, "action": "allow" },    { "detectorId": "prompt.unicode_obfuscation", "score": 0, "action": "allow" },    { "detectorId": "text.pii", "score": 1, "action": "redact" }  ],  "text": "my email is [redacted email]",  "latencyMs": 0.05,  "degraded": false}
ActionMeaningWhat the OpenAI wrapper does
allowNothing firedThe text passes through unchanged
flagSomething looks suspicious but is not certainThe turn continues and the verdict is recorded for review
redactSensitive spans were found and replacedThe turn continues with the redacted text
blockThe text must not get throughThe call raises a blocked error

Every detector that runs reports its own action and a score between 0 and 1. The verdict takes the most severe action among them, in the order allow, flag, redact, block. When two detectors pick the same action, the higher score wins. The detectors array keeps every outcome, so you can always see why a verdict came out the way it did.

Detectors

A detector looks for one kind of problem. Each one runs in a fixed phase and has a default action:

  • Rule-based detectors match concrete patterns such as an email address, an API key or your own system prompt. They run in microseconds and give the same answer every time
  • The injection classifier (prompt.injection_classifier) is a model that scores prompts for injection and jailbreak attempts. It runs on input only, and it flags rather than blocks
  • The judge is a language model that reviews verdicts the cheaper layers are unsure about. It never overrules a redact or block from a rule-based detector, because those are matches, not guesses

See Guardrails for every detector, its phase and what it catches.

Profiles

A profile picks the detection layers for a check and sets its time budget. Pass it with each check. When you leave it out, Verexa uses deterministic.

ProfileBudgetLayersJudge
deterministic30 msSix rule-based detectorsNever
balanced1.5 sRule-based detectors and the injection classifierIn the background, for uncertain verdicts
audit8 sRule-based detectors and the injection classifierBefore answering, for every verdict that is not already certain

With balanced, the judge runs after Verexa has already answered. Its verdict does not change the response you got, and it lands in the same trace a moment later. With audit, the response waits for the judge, so the judge's verdict counts.

Check a prompt with the audit profile
from verexa import create_guardguard = create_guard(timeout_ms=15000)verdict = guard.check_input(user_message, profile="audit")
Raise the timeout for audit
The SDKs give up on a check after 2 seconds by default. That is enough for deterministic and balanced, but an audit check can take up to 8 seconds. Raise the timeout to 15 seconds for audit checks, or the SDK abandons the check and fails open before the verdict arrives.

The OpenAI wrappers always use the default profile. To use balanced or audit, call the guard directly.

Policies and monitor mode

A policy adjusts a profile for one project and environment. The API key you send decides which project and environment a check belongs to, so the right policy applies without any extra fields. For each detector in the profile, a policy can:

  • Disable it so it does not run at all
  • Put it in monitor mode so it still runs and is recorded, but its action is capped at flag. A monitored detector never blocks or redacts

Monitor mode is the safe way to try a detector on real traffic. Watch what it would have caught in the dashboard, then switch it to enforce. See Policies for how to set one up.

Traces

A trace id groups every check that belongs to one conversation turn. The OpenAI wrappers create one per call and use it for both the input and the output check. When you call the guard yourself, create one with randomTraceId() (random_trace_id() in Python, verexa.NewTraceID() in Go) and pass it to both checks.

In the dashboard, Events shows each trace as one row with its checks, the detectors that fired and the action taken. A background judge verdict from the balanced profile is added to the same trace when it finishes.

Degraded verdicts

A verdict has degraded: true when Verexa could not run everything the profile asked for. That happens when the budget runs out before every detector has run, or when the classifier or the judge cannot be reached. Detectors that were skipped are listed in degradedDetectors.

A degraded verdict is still a real verdict, built from the detectors that did run. Verexa does not cache it, so the next check of the same text tries the full set of layers again.

Fail open and fail closed

If the SDK cannot get a verdict at all, it falls back to one of its own. That covers a missing API key, a network error, a timeout and a 5xx from the API. The fallback has degraded: true and an action that depends on the fail mode:

  • Fail open (the default) returns allow, so your app keeps working while Verexa is unreachable
  • Fail closed returns block, so no text gets through unchecked
Fail closed
from verexa import create_guardguard = create_guard(fail_mode="closed")

After 5 failed checks in a row, the SDK stops calling the API for 30 seconds and returns the fallback straight away, so an outage does not add a timeout to every request. A 4xx response does not count toward this, because it means your key or request is wrong, not that the API is down.

A fallback allow looks like a clean verdict to code that only reads action. If you need to know whether a check really ran, read degraded too. See Handle failures.

Next steps

  • Guardrails: every detector and what it catches
  • Policies: turn detectors on or off and run them in monitor mode
  • Handle failures: react to blocked turns, degraded verdicts and outages
  • Check API: every request and response field