Skip to main content
Guardrails apply local input checks before a request reaches the model provider. Use them to block recognized prompt-injection and jailbreak patterns, and to block or redact sensitive data. They apply to application, integration, and Playground model requests under the same project.

Configure a policy

  1. Select All projects for an organization policy, or select a project for a project policy.
  2. Open Guardrails and create a policy with a name and description.
  3. Choose Prompt injection, Jailbreak attempts, and/or Sensitive data.
  4. Select the sensitivity. For sensitive data, choose the individual data kinds and Block or Redact for each rule.
  5. Test representative samples with the draft tester, review the result, and save the enabled policy.
Organization policies apply across projects. Project policies add protection; they cannot weaken an organization policy. Project admins can change their project policies, while organization administrators manage shared policies.

Choose protections

For Company email, enter the protected domains; subdomains are included. Custom patterns use RE2, which supports predictable matching without backreferences or lookaround. A policy accepts up to 20 custom patterns, each at most 200 characters, and up to 50 protected email domains.

Block or redact

Block rejects the request before an upstream connection is opened. The error identifies the policy and detector without echoing the matched text. Redact replaces matches in the prompt fields sent upstream. The draft tester lets you inspect the transformed model input before enabling the rule. Test both content that should match and ordinary content that should pass, especially when using strict sensitivity or custom patterns.

Verify and operate a policy

Use the tester first, then send representative requests in Playground. A pattern detector can miss unfamiliar attacks or match legitimate text; keep application access controls and review agent actions independently. Guardrails inspect incoming content, so assess model output in the application that uses it. The tester evaluates samples in memory without saving them. Request-level diagnostic capture is an independent, explicit Activity setting.