AIOBES logo header

AIOBES Research Catalogue

This catalogue links to research papers relevant to the AIOBES standard. These works define behavioural evaluation methods, testing protocols and analysis frameworks that support the development and application of AIOBES.

AIOBES Foundational Paper

The AIOBES Foundational Paper defines the core structure, purpose and operational expectations of the Artificial Intelligence Operational Behaviour Evaluation Standards. It establishes the baseline behavioural governance model that all future AIOBES documents build upon.

The paper outlines the foundational principles, scope and evaluative requirements needed to support consistent behavioural assurance across AI systems. It acts as the primary reference point for the development, alignment and evolution of the AIOBES standard.

Zenodo DOI

Vectored Conversational AI Testing

Vectored Conversational AI Testing is a behavioural evaluation method for AI systems that operate through conversation. It uses live interaction to observe coherence, boundary handling, context retention and behavioural stability as dialogue evolves.

Controlled conversational variations reveal behavioural patterns that static tests do not expose. The paper outlines the structure and operational workflow of the method, its relevance to EU AI Act expectations and the boundaries of what the approach does and does not attempt.

Zenodo DOI

LLM INQUISITOR Methodology (GitHub Edition) v1.1

This document defines the LLM INQUISITOR Methodology: a structured, repeatable discipline for evaluating the behaviour of large language models under controlled load. It provides a formal approach for assessing reliability through observable behaviour and evidentiary traceability.

The methodology supports rigorous evaluation in research, safety and enterprise assurance contexts, where behavioural stability under real operational conditions is essential.

Zenodo DOI

Argo AI Testing Protocol: Sustained Multi Axis Load Testing

Most evaluation of conversational AI relies on short, prompt based tests that fail to reflect how real people use these systems. Such tests do not capture extended interaction, shifting user intent or cumulative context effects.

This paper introduces the Argo AI Testing Protocol, a conceptual approach for evaluating AI systems within the User Interaction Space, the full set of observable outputs and interactions available to a user.

Zenodo DOI

Argo’s Fundamentals of Failings in Prompt Test Design and Evaluation for LLMs

This paper identifies the core structural failings in prompt test design and evaluation for LLMs. It shows that current methods cannot produce reliable signals: they mismeasure capability, misinterpret outputs and often generate failure states created by the tests themselves.

These practices emerged in an industry expanding faster than it can define standards, leaving evaluation shaped by inconsistent methods and gatekeepers with limited grounding in the systems they are judging.

Zenodo DOI