Inquisitor Labs
About Inquisitor Labs

Inquisitor Labs develops behavioural testing methodologies, visibility audits, workflow tools, UX evaluations, operational standards and research that show how AI actually behaves when real people and real work interact with it. Our focus is practical, evidence-driven evaluation -exposing failure modes, misinformation and workflow-level risks that prompt-based testing will never surface.

AI systems fail in deployment because their behaviour isn’t understood, measured or tested under realistic conditions. That’s the gap we specialise in. Everything we publish and every service we offer exists for one reason - to give organisations clear behavioural evidence, stronger operational control, and a realistic understanding of how their AI systems perform in the real world.

Testing Icon

AI Brand Visibility Audit Service

AI is talking about your brand behind your back.

Customers ask chatbots if you’re safe, legitimate, worth the money, or better than competitors - and AI answers with whatever narrative it has absorbed.

ChatGPT alone now answers billions of questions each day. Some of those questions are about your brand.

An AI Brand Visibility Audit gives you the missing layer now standing between your brand and the public.
Find out what AI is saying.

AI & UX Evaluation Icon

AI & UX Evaluation Service

We evaluate AI systems using real user interactions across real contexts – from everyday work tasks to AI‑driven apps and chatbots. We help identify AI & UX failures, behavioural issues and points of friction early, before they turn into customer impact or regulatory and financial consequences.

EU Flag

EU AI Act – Support, Resources and Guidance

Support, resources and guidance for organisations that must comply with the EU AI Act. Find out whether the Act applies to your systems, understand what evidence is required, and learn how to test AI using real‑world conditions to meet the new legal obligations.

AIOBES Image

AIOBES – Behavioural Governance Standard

The emerging standard for AI behavioural safety, reliability, and structured evaluation across modern AI systems.

Working with AI Icon

How to Work Effectively with AI

Let’s face it - AI does not always live up to the hype. It goes off‑track, changes tone, forgets things, and can give different (or inaccurate) answers to the same request. Many people don’t know why it happens or how to make it more manageable. So we’ve developed some lightweight tools and straightforward tips to help you get more out of AI in your everyday work.

Archive Image

Archive of AI Anomalies

Strange and unusual behaviour from the frontiers of latent space. When we push the edges of AI behaviour in our reseach, we often get surprising results. Some have been declassified here as a public resource.

Research Image

Published Research

Public research linked to Inquisitor Labs. Published on Zendo. Updates will be added as further work is completed.

Testing Icon

AI Behavioral Testing Under Realistic Conditions Matters

AI systems need to be evaluated using realistic conditions, behavioural evaluation, and scenario‑based analysis. It is the only way to identify failures, measure reliability, and generate evidence that aligns with modern AI governance and regulatory requirements. Find out more here.

LLM Inquisitor AI Testing Methodology

LLM Inquisitor Cover

Convesrational AI Testing

Vectored Conversational AI Testing

“There is no magic prompt with which you can test AI. What you can do, is give AI real work under realistic conditions and observe what happens.”

LLM Inquisitor is a practical methodology for evaluating and stress-testing AI systems under real-world conditions. It exposes hidden failure modes, behavioural drift, data leakage, hallucinations, and workflow-level risks that traditional prompt testing cannot detect. Structured, scenario-based, and auditable, it gives developers, testers, and governance teams a repeatable way to understand and manage AI behaviour.

Open on Leanpub

“Users do not ask one question and leave. They engage in multi-turn, free-flowing dialogue.”

Vectored Conversational AI Testing is a behavioural methodology built for real conversational interactions, not single-prompt checks. It evaluates how AI systems adapt, drift, recover, escalate, or fail across full conversational arcs. Structured, repeatable, auditable -designed for developers, governance teams, and standards bodies who need reliable behavioural evidence.

Open on Leanpub

Medium Logo

Articles Published On Medium

Read our articles on AI & Tech published on Medium.

Reddit Image

Join the discussion on Reddit

Read about and discuss AI evaluation topics, methods and resources on r/llm_ai_eval.

(C) William Argo

Contact email: inquisitor.labs@proton.me