Inquisitor Labs
Testing Icon

AI Behavioural Testing

AI systems do not fail at the level of prompts. They fail at the level of behaviour. Behavioural testing examines how a system behaves across time, across workflows, and across realistic conversational or operational sequences.

It focuses on stability, predictability, behavioural integrity, and real workflow performance. This page introduces the core ideas behind AI behavioural testing and links to the pages that expand each part of the discipline.

Core Concepts

Why Behavioural Evidence Matters

Legal teams explain what the law demands, but not whether you have the evidence. Auditors confirm whether evidence exists, but not how to produce it. Behavioural testing fills this gap by showing how to generate behavioural evidence suitable for real deployments.

It examines how systems behave across time and workflows, not just in isolated prompt tests, so you can stay compliant and avoid liability.

Why Prompt Testing Is Not Enough

Prompt testing treats AI systems as if they were simple input output machines. Modern models do not behave this way. They change state, accumulate context, shift behaviour as interactions progress, and behave differently under pressure or load.

Prompt testing cannot reveal whether a system is stable, predictable, or safe across time. It cannot show drift, collapse, or behavioural divergence. These failures only appear when you test behaviour, not prompts.

See Prompt Testing.

Behavioural Failure Patterns

Real deployments reveal consistent behavioural patterns that never appear in single prompt tests. These patterns are central to understanding system stability and reliability.

  • behaviour changing over time
  • behaviour diverging across similar inputs
  • behaviour collapsing under pressure

These patterns are invisible to audits, dashboards, and compliance checklists. They only emerge when you observe the system in motion, not in snapshots.

See Behavioural Failure Modes.

Behavioural Testing vs Prompt Testing

Behavioural testing and prompt testing measure different things. Understanding the distinction is essential for evaluating system quality, assessing deployment risk, and avoiding personal and corporate liability.

  • evaluating system quality
  • assessing deployment risk
  • avoiding personal and corporate liability

Prompt testing tells you nothing about whether a system will behave safely tomorrow. Behavioural testing shows whether the system is stable enough to trust and whether you have the evidence required for legal and operational accountability.

See Behavioural Testing vs Prompt Testing.

Vectored Conversations

Vectored conversations provide a structured way to observe system behaviour across controlled conversational paths. They support repeatable behavioural observation, workflow simulation, stability and drift detection, and cross model comparison.

  • repeatable behavioural observation
  • workflow simulation
  • stability and drift detection
  • cross model comparison

See Vectored Conversations.

LLM INQUISITOR Methodology

The LLM INQUISITOR Methodology is the first behavioural testing methodology designed specifically for LLM systems. It provides a complete, structured approach for evaluating AI behaviour in real workflows and under real conditions.

It is the only approach designed to produce behavioural evidence suitable for legal, operational, and compliance use.

The Methodology page contains the full paper, including:

  • the behavioural testing lifecycle
  • scenario and workflow design
  • behavioural signature extraction
  • evaluation patterns
  • practitioner guidance

See Methodology.