Inquisitor Labs

Published Research

Publicly available research linked to Inquisitor Labs can be found here. This is not an exhaustive list. Updates will be posted as further research is completed.

Research Catalogue

AIOBES Foundational Paper

The Artificial Intelligence Operational Behavioural Evaluation Standards (AIOBES) Foundational Paper defines the core behavioural governance principles for AI systems, establishing the baseline structure for the AIOBES standard. It outlines the foundational concepts, scope, and operational expectations required for consistent behavioural evaluation and governance across AI systems.

The paper provides the initial framework that future AIOBES documents will build upon, forming the structural basis for behavioural assurance, evaluative consistency, and standardised governance practices.

Zenodo DOI

Vectored Conversational AI Testing

Vectored Conversational AI Testing is a behavioural evaluation method for AI systems that operate through conversation. It uses live interaction to observe coherence, boundary handling, context retention, and behavioural stability as dialogue evolves.

Controlled conversational variations reveal behavioural patterns that static tests do not expose. The paper outlines the structure and operational workflow of the method, its relevance to EU AI Act expectations, and the boundaries of what the approach does and does not attempt.

Zenodo DOI

LLM INQUISITOR Methodology (GitHub Edition) v1.1

This document defines the LLM INQUISITOR Methodology: a structured, repeatable discipline for evaluating the behaviour of large language models under controlled load. It provides a formal approach for assessing reliability through observable behaviour and evidentiary traceability.

The methodology supports rigorous evaluation in research, safety, and enterprise assurance contexts, where behavioural stability under real operational conditions is essential.

Zenodo DOI

Argo AI Testing Protocol: Sustained Multi Axis Load Testing

Most evaluation of conversational AI relies on short, prompt based tests that fail to reflect how real people use these systems. Such tests do not capture extended interaction, shifting user intent, or cumulative context effects.

This paper introduces the Argo AI Testing Protocol, a conceptual approach for evaluating AI systems within the User Interaction Space — the full set of observable outputs and interactions available to a user.

Zenodo DOI

Argo’s Fundamentals of Failings in Prompt Test Design and Evaluation for LLMs

This paper identifies the core structural failings in prompt test design and evaluation for LLMs. It shows that current methods cannot produce reliable signals: they mismeasure capability, misinterpret outputs, and often generate failure states created by the tests themselves.

These practices emerged in an industry expanding faster than it can define standards, leaving evaluation shaped by inconsistent methods and gatekeepers with limited grounding in the systems they are judging.

Zenodo DOI

Argo Prompting: Pattern Formation in LLMs Under Sustained Conceptual Pressure

This paper introduces Argo Prompting, a method for inducing pattern formation behaviour in large language models through sustained conceptual pressure. It distinguishes pattern formation from collapse, hallucination, and surface level pattern matching.

The paper provides a practical framework for researchers studying LLM behaviour under extended reasoning conditions.

Zenodo DOI