AIOBES logo header

AIOBES

Artificial Intelligence Operational Behaviour Evaluation Standards

AIOBES is the only testing standard and the only set of methodologies that satisfy EU AI Act behavioural evaluation requirements.

AIOBES is the first comprehensive behavioural evaluation standard for Artificial Intelligence systems. It focuses on how AI actually behaves when people use it – in real tasks, real workflows and real conversations.

AIOBES evaluates operational behaviour rather than isolated prompt testing. It identifies behavioural failures that create regulatory, legal, operational and reputational exposure, and provides a structured way to detect, document and evidence these behaviours.

The AIOBES Foundational Paper defines the behavioural domain, evidentiary structure and operational methodologies that make AIOBES a governance‑grade standard. This section also provides resources for standards bodies and regulators seeking to incorporate AIOBES into external standards, governance frameworks or regulatory instruments. You can read the foundational paper or access the formal integration specification below.

AIOBES Developer Adoption Guide (PDF) AIOBES Developer Adoption Guide (Page)
AIOBES Corporate Adoption (PDF) AIOBES Corporate Adoption (Page)
AIOBES Standards Body Guide (PDF) AIOBES Standards Body Guide (Page)
Read the AIOBES Foundational Paper View and Download the PDF on Zenodo

AIOBES Scope

AIOBES evaluates the operational behaviour of AI systems used in real tasks, workflows and interactions. It applies to general purpose AI systems, conversational AI, agentic systems, workflow automation, multimodal systems and domain specific AI used in regulated environments.

AIOBES does not evaluate model architecture, training data, internal mechanisms or synthetic benchmark performance. It evaluates only operational behaviour: how the system behaves when people use it.

AIOBES Assessment Dimensions

AIOBES evaluates AI systems across behavioural dimensions that define what must be examined during behavioural testing. These dimensions identify the areas where stability, instability and contamination typically occur.

  • reproducible behaviour or reproducible failure patterns
  • stable or unstable outputs under repeated inputs
  • correct alignment or misalignment with task requirements
  • consistent or inconsistent guardrail behaviour
  • presence or absence of workflow contamination
  • clear, ambiguous or undocumented failure modes
  • behavioural stability or behavioural instability across real interactions

AIOBES is designed to detect and document behavioural issues clearly and without ambiguity whenever they occur.

AIOBES Standard Components

The AIOBES standard is built from six core components. Each provides a specific structural function within the overall framework.

These components form the structural backbone of AIOBES and provide the foundation for behavioural evaluation, documentation and compliance.

What AIOBES Evaluates

AIOBES focuses on behavioural failure modes that appear under real use such as:

  • context collapse
  • misaligned or inconsistent guardrails
  • hallucinated content
  • contradictory or unstable outputs
  • incorrect refusals
  • misaligned responses in regulated workflows
  • tone instability

These failure modes directly cause workflow contamination - where incorrect, unsafe or non-compliant outputs spread through processes and systems.

Operational Effects Exposed by AIOBES

When these failure modes occur, organisations face:

  • workflow collapse
  • workflow contamination
  • regulatory non-compliance
  • fines and sanctions
  • legal liability exposure
  • operational disruption
  • reputational damage

AIOBES documents these effects clearly, repeatably and without ambiguity.

AIOBES Evaluation Criteria

AIOBES uses five behavioural criteria to assess how an AI system performs under real operational conditions. These criteria define the behavioural areas where stability, instability, alignment, misalignment and contamination are observed and documented.

  • Stability - consistent or inconsistent outputs under identical conditions
  • Alignment - adherence or drift from task requirements
  • Integrity - presence or absence of hallucination, fabrication or contamination
  • Continuity - maintenance or collapse of structure across long-form work
  • Coherence - conversational correctness or instability across varied interactions

AIOBES Evidence Categories

AIOBES uses three evidence categories:

  • Input Evidence - prompts, documents, interaction patterns and multimodal triggers
  • Output Evidence - all raw outputs including correct outputs, incorrect outputs, instability, drift, collapse, contamination and guardrail behaviour
  • Operational Evidence - workflow impact, contamination, collapse, regulatory exposure and downstream consequences

Operational Concepts

Operational concepts within AIOBES describe the conditions under which behavioural evaluation takes place. They outline the operational environment an AI system is expected to perform in and the behavioural boundaries it must maintain across variations in user intent, context and load.

These concepts identify the areas that form the foundation of operational behavioural evaluation. They include behavioural stability, operational boundaries, drift behaviour, load conditions, interaction space and the generation of operational behavioural evidence. These areas will be developed into the full operational evaluation model.

AIOBES will expand these operational concepts into a dedicated operational section. This will define the complete operational framework used to assess AI behaviour under real or realistically simulated conditions.

AIOBES Evaluation Stack

  • Risk (as defined by the EU AI Act)
  • Input (comprehensive, realistic)
  • Behavioural Observation
  • Failure Mode Analysis
  • Evidence Packaging

AIOBES Compliance Pathway

AIOBES compliance is demonstrated through captured evidence showing how the system behaves under real operational conditions.

  • Identify the system’s operational role and classification under current EU AI Act standards
  • Run behavioural evaluation using AIOBES methodologies
  • Capture input evidence, output evidence and operational evidence
  • Document behavioural stability, alignment, continuity, coherence and integrity
  • Publish an AIOBES Compliance Statement

AIOBES and the EU AI Act

AIOBES provides behavioural evidence required under multiple EU AI Act obligations including transparency, risk classification, behavioural stability, misuse prevention, safety documentation and post-market monitoring.

AIOBES does not replace legal compliance. It provides the behavioural evidence required to demonstrate it.

Methodologies

AIOBES uses two operational methodologies to expose how AI behaves under real conditions. They show how specific failure modes lead directly to workflow contamination, workflow collapse and the resulting regulatory, legal, operational and reputational exposure.

LLM Inquisitor

Evaluates how AI behaves during long-form work using real tasks, real documents and real operational content. It exposes long-form collapse where the AI loses structure, drifts from requirements, rewrites instead of editing, contaminates the work or breaks the task over time.

It shows whether the AI can maintain stability, alignment and continuity across extended, multi-step real-world work without introducing errors, hallucinations or workflow contamination.

Vectored Conversational AI Testing

Evaluates AI behaviour across real conversational interactions - text, verbal and multimodal. It examines how the system responds under varied interaction patterns, phrasing, formats and multimodal triggers.

It reveals whether the AI maintains coherence, correctness and alignment when engaged through natural user communication, and whether conversational instability contaminates downstream workflows.

Reference Materials

Evaluating and Testing AI Using Real World Work

Covers long-form work collapse, real task evaluation and contamination detection in extended operational work.

Available here: https://leanpub.com/llm-inquisitor

Vectored Conversational AI Testing Guide

Covers behavioural evaluation across real conversational interactions and how conversational instability contaminates downstream workflows.

Available here: coming soon to Leanpub

These guides form the baseline for applying AIOBES correctly.

Training (Coming Soon)

Practical training for applying AIOBES in real workflows and real interactions. Focused on operational execution.

Certification (Coming Soon)

Formal practitioner certification aligned to the two methodologies. Validates correct application of AIOBES and production of audit-ready behavioural evidence.

Compliance (Q4 2026)

Organisational compliance framework based on AIOBES. Provides a structured way to demonstrate behavioural evaluation controls to regulators.