Publicly available research linked to Inquisitor Labs can be found here. This is not an exhaustive list. Updates will be posted as further research is completed.
This paper presents the Criterion Audit, the leading‑edge methodology for evaluating how AI systems surface, frame, suppress, or distort brand visibility during user‑initiated discovery. It replaces generic AI‑ethics checklists and inconsistent prompt‑testing with a fixed behavioural taxonomy, stable scoring anchors, and a completeness check that prevents partial or invalid audits.
It is the only framework built specifically for brand‑visibility mediation rather than broad AI quality assessment, eliminating the ambiguity and subjectivity found in existing approaches. The methodology is published openly; the proprietary question battery and execution heuristics remain restricted.
This paper examines how conversational AI systems mediate brand visibility during user‑initiated discovery. It explains why certain brands are surfaced more readily than others, how structural behaviours in AI interaction create representational imbalance, and what the implications are for organisations.
The paper highlights the underlying mechanisms that distort exposure and outlines why evaluating AI‑mediated discovery pathways is necessary, without prescribing a specific audit method.
The Artificial Intelligence Operational Behavioural Evaluation Standards (AIOBES) Foundational Paper defines the core behavioural governance principles for AI systems, establishing the baseline structure for the AIOBES standard. It outlines the foundational concepts, scope, and operational expectations required for consistent behavioural evaluation and governance across AI systems.
The paper provides the initial framework that future AIOBES documents will build upon, forming the structural basis for behavioural assurance, evaluative consistency, and standardised governance practices.
Vectored Conversational AI Testing is a behavioural evaluation method for AI systems that operate through conversation. It uses live interaction to observe coherence, boundary handling, context retention, and behavioural stability as dialogue evolves.
Controlled conversational variations reveal behavioural patterns that static tests do not expose. The paper outlines the structure and operational workflow of the method, its relevance to EU AI Act expectations, and the boundaries of what the approach does and does not attempt.
This document defines the LLM INQUISITOR Methodology: a structured, repeatable discipline for evaluating the behaviour of large language models under controlled load. It provides a formal approach for assessing reliability through observable behaviour and evidentiary traceability.
The methodology supports rigorous evaluation in research, safety, and enterprise assurance contexts, where behavioural stability under real operational conditions is essential.
Most evaluation of conversational AI relies on short, prompt based tests that fail to reflect how real people use these systems. Such tests do not capture extended interaction, shifting user intent, or cumulative context effects.
This paper introduces the Argo AI Testing Protocol, a conceptual approach for evaluating AI systems within the User Interaction Space — the full set of observable outputs and interactions available to a user.
This paper identifies the core structural failings in prompt test design and evaluation for LLMs. It shows that current methods cannot produce reliable signals: they mismeasure capability, misinterpret outputs, and often generate failure states created by the tests themselves.
These practices emerged in an industry expanding faster than it can define standards, leaving evaluation shaped by inconsistent methods and gatekeepers with limited grounding in the systems they are judging.
This paper introduces Argo Prompting, a method for inducing pattern formation behaviour in large language models through sustained conceptual pressure. It distinguishes pattern formation from collapse, hallucination, and surface level pattern matching.
The paper provides a practical framework for researchers studying LLM behaviour under extended reasoning conditions.