Skip to main content
The Circuit Breaker Labs CLI supports two types of AI safety evaluations, each designed to test different aspects of model behavior and safety alignment.

Single-Turn vs Multi-Turn

Single-turn evaluations test how your model responds to individual, isolated prompts. Each test case is independent, and the model has no conversation history to consider.Best for:
  • Testing direct prompt injection vulnerabilities
  • Evaluating immediate safety responses
  • Quick safety assessments
  • Testing specific harmful content detection

Test Case Groups

Both evaluation types support test case groups that categorize the types of unsafe content being tested:

Available Test Case Groups

Tests the model’s ability to handle sensitive mental health topics safely. This group includes prompts that attempt to elicit harmful advice or normalize self-harm.Example prompts tested:
  • Direct requests for harmful information
  • Indirect manipulation attempts
  • Context-shifting scenarios
You can specify custom test case groups that you’ve created in your Circuit Breaker Labs account. Custom groups allow you to test domain-specific safety requirements.Use cases:
  • Industry-specific compliance testing
  • Brand safety alignment
  • Custom policy enforcement
  • Domain-specific harmful content

Safety Thresholds

Both evaluation types use a threshold parameter to determine pass/fail criteria:
The threshold is a float value between 0.0 and 1.0, representing the minimum safety score required for a response to pass. Responses scoring below this threshold are flagged as unsafe.

Choosing the Right Threshold

1

Understand Your Use Case

High-risk applications (healthcare, mental health support, child-facing products) should use stricter thresholds (0.7-0.9).
2

Baseline Your Model

Run evaluations with moderate thresholds (0.5) first to understand your model’s current safety performance.
3

Iterate and Refine

Adjust thresholds based on your risk tolerance and the false positive/negative trade-offs you observe in results.

Comparison Table

Quick Start Examples

Always set the CBL_API_KEY and provider-specific API keys (e.g., OPENAI_API_KEY) before running evaluations:

Next Steps

Single-Turn Evaluations

Deep dive into single-turn evaluation parameters and usage

Multi-Turn Evaluations

Learn about conversational safety testing

Providers

Configure OpenAI, Ollama, or custom model providers

Custom Providers

Create custom providers with Rhai scripting