Documentation
Open app

Guardrails

On this page

Guardrails keep conversations safe: they block prompt-injection and jailbreak attempts, mask personal data, and screen tool/document output. Configure them from the workspace's direct Guardrails navigation item. Available to authorized organization owners/admins.

What guardrails check#

The Hub exposes two workspace-controlled planes. Core also applies a platform-managed reasoning safeguard:

What is checked
Mode setting in the HubMeaning
User input enforcementYour messagesThe text people type, checked before the assistant answers. Personal data (emails, phone numbers) can also be masked here.
Tool output enforcementTool & document contentText the assistant brings in from tools, uploaded files, the web, or the knowledge base — checked before it is used.
Model reasoningReasoning on its way to the browser or storagePlatform-managed detection masks secrets and personal data before reasoning is streamed or saved. This plane is not a workspace mode selector.

What each mode does#

You choose a mode separately for User input enforcement and Tool output enforcement. Here is what each mode does in both cases:

ModeUser input enforcementTool output enforcement
OffWorkspace-selected input checks do not run.Workspace-selected tool-output checks do not run.
Observe (log only)A match is recorded in the log; nothing is changed or blocked.A match is recorded in the log; nothing is changed or blocked.
Redact (mask)Personal data (e.g. an email) is hidden before the assistant sees it.The matched part is hidden when possible, then the content continues.
Enforce (block)A risky message (e.g. a jailbreak attempt) is blocked; personal data is still masked.Unsafe content is normally withheld and a safety notice is shown. Knowledge Base app tool results are capped at soft enforcement: matches can be logged or redacted but the retrieved document is not blocked.
Platform protections still apply

The workspace master switch and Off disable workspace-selected checks only. Platform-mandatory guards can still run, and the Hub marks them as Enforced by the platform.

Turn guardrails on for a workspace#

1

Open the Guardrail Hub

Open the workspace in Console and choose Guardrails from its sidebar. If it is absent, your current role cannot view it.

The Guardrails tab showing platform-enforced safeguards, model-reasoning masks, User input set to Enforce, Tool result set to Redact, and two of eleven workspace guards selected
The Hub separates platform-enforced safeguards from workspace controls; this example enforces user input, redacts tool results, and selects 2 of 11 guards.
2

Turn on the master switch

Toggle Enable guardrails for this workspace. This controls workspace-selected guards; platform-mandatory protections remain active.

3

Choose how strict to be

Set how guardrails act on your messages and on tool output (Observe, Redact, or Enforce). A short note under each option explains what it does.

4

Pick which guards run

Tick the guards you want for this workspace; search or filter to find them. Leave the list untouched to keep the recommended defaults. A few guards need some values from you (a list of banned words, competitor names, or allowed topics) — fill those in, otherwise you can't save.

5

Save

Click Save. Changes take effect immediately for new messages. Use Export rules (CSV) to download the current configuration.

Available guards#

Guards marked Recommended are on by default. For the last three, you provide the list they check against.

The recent blocks log listing blocked, redacted, shadow, and model-reasoning events with their stage, guard, and a masked text preview
Recent blocks records input, tool-result, and model-reasoning detections without exposing the sensitive text.
GuardProtects againstYou provide
Jailbreak / Prompt InjectionAttempts to trick or hijack the assistant, including hidden instructions inside documentsNothing (Recommended)
Secrets & CredentialsAPI keys, passwords and tokens in a messageNothing (Recommended)
Personal Data (PII)Emails, phone numbers and other personal data (masked)Nothing (Recommended)
Toxic LanguageToxic, hateful or harassing languageNothing
ProfanityProfane or obscene wordsNothing
NSFW TextSexually explicit / not-safe-for-work textNothing
Toxic Language (Multilingual)Toxic language across languages, including Vietnamese and JapaneseNothing
Gibberish / NonsenseGarbled or nonsensical inputNothing
Unusual / Manipulative PromptManipulative or social-engineering promptsNothing
Banned WordsWords or phrases you don't allowYour list of words
Competitor MentionsMentions of competitors you nameCompetitor names
Restrict to TopicsAnything outside the topics you allowAllowed topics

Who can view and edit guardrails#

Guardrails use a single permission: whoever can open the Hub can also change it. There is no read-only view — the configuration, the guard list, and the Recent blocks log are all covered by the same right to manage the workspace.

RoleGuardrails access
Organization OwnerYes — in every workspace of the organization
Organization AdminYes — in every workspace of the organization
Workspace AdminYes — in the workspaces they administer
Workspace MemberNo — the Guardrails item is not shown
Organization Member with no workspace admin roleNo

Organization Owners and Admins receive workspace-admin rights in every workspace automatically, so granting either role hands over guardrail control for the whole organization. To limit someone to one workspace, leave them an Organization Member and make them a Workspace Admin there. SotaAgents platform support staff can also reach the Hub when assisting you.

Hiding the menu is not the control

The missing sidebar item only reflects the same rule the API enforces: a request from an account without workspace-admin rights is rejected, so guardrails cannot be read or changed by other means.

Which guards the platform enforces#

Two guards in the table above are platform-mandatory by default. They run on every workspace even when the master switch is off or a mode is set to Off, and they have no tickbox:

GuardWho controls it
Jailbreak / Prompt InjectionPlatform — always on, including the screening of tool and document output
Secrets & CredentialsPlatform — always on
Personal Data (PII)You — except in model reasoning, where the platform masks it
The other nine guardsYou — tick, untick, and choose the mode per workspace

Model reasoning is also platform-managed: secrets and personal data are masked there regardless of which guards you tick. A platform-mandatory row follows the platform's own mode, not your workspace mode, so it can be blocking while your workspace sits on Observe.

Your deployment is the final word on this, because the mandatory set is an operator setting. Read it off the Hub: everything under Enforced by the platform is out of your hands, and everything with a tickbox is yours. If that section is absent, the platform layer is switched off in your deployment and all twelve guards are workspace-controlled.

How long guardrail logs are kept#

Recent blocks keeps events for 30 days by default, after which each entry is deleted automatically. The window is a platform setting, so a self-hosted or dedicated deployment may be configured differently — ask your platform operator if you need the exact number for an audit.

Each entry records the time, the stage (user input, tool output, or model reasoning), what happened (blocked, redacted, observed, or not evaluated), the guard that fired, and the mode in force. The flagged message itself is never stored: entries keep its length and a short fingerprint, plus a masked excerpt of at most 280 characters. For secret and personal-data findings even that excerpt is omitted, so the log cannot leak what the guard just caught.

Treat the log as a debugging aid, not an archive

It answers "which message was stopped, where, and by which guard" while the incident is recent. If you must retain guardrail activity for longer than the retention window, copy what you need out before it expires. Export rules (CSV) exports the configuration, not the log.

Contents

Esc

Search titles and body text across every chapter.