Glossary
This glossary collects all terminology defined and coined across the architectural
specifications and methodologies published by Inquisitor Labs.
Terms are grouped by paper, in publication date order, and each section contains
the architectural definitions exactly as they appear in the original documents.
This page serves as the canonical reference index for all system‑level concepts, behaviours,
classifications and architectural constructs used throughout the Inquisitor Labs research canon.
1. Argo Prompting: Pattern‑Formation in LLMs Under Sustained Conceptual Pressure
Argo Prompt / Argo Prompting
The generic term for the conversational style in which the user pushes a model into
unfamiliar conceptual territory while keeping it coherent through sustained constraint. It
refers to the interaction itself -the way the user introduces new structure, maintains the
frame, and prevents drift- rather than to any internal mechanism or special capability.
Argo Prompting is simply the name for this style of concept-pushing dialogue.
Behavioural regime
A recognisable mode of model behaviour that persists across several turns when the
conversation is held within a stable frame. Used descriptively to distinguish ordinary
interpolation from the more coherent behaviour that sometimes appears under
sustained conceptual pressure.
Conceptual pressure
The effect of keeping the conversation inside an unfamiliar conceptual region while
preventing the model from drifting back to generic or safer patterns. It is a property of the
interaction, not a claim about internal states.
Constrained frame
A stable set of expectations, boundaries, and prohibitions maintained by the user across
an extended interaction. It prevents drift and allows more demanding behaviour to
emerge.
Constraint (final constraint)
The last structural piece supplied by the user that completes the pattern the model has
been circling. Introducing it often triggers the abrupt coherence shift associated with the
OMG moment.
Coherence shift
A noticeable increase in clarity, structure, or internal consistency in the model’s output.
It is an observable change in surface behaviour, not evidence of internal recognition.
Completed pattern
The temporarily stabilised structure that becomes available after the coherence shift.
Once present, the model can reason within it more consistently until the structure
decays.
Exclamatory spike
A brief, surface-level outburst -sometimes emphatic or profane- that marks the moment
of abrupt coherence increase. It is a behavioural marker, not an emotional reaction.
Extended-pattern
A temporarily stabilised region of coherent behaviour that appears after the OMG
moment. It persists only while the conversational constraints remain in place.
Pattern
A locally coherent behaviour the model is already producing. It refers to the structure
present in the interaction at that moment, not to any stored representation.
Pattern-extension
When the model continues or stretches an existing pattern because the interaction keeps
it moving in the same conceptual direction. It is the model staying with the structure it
already has, not creating a new one.
Pattern-formation
When a new, recognisable pattern appears and stabilises enough to be used in
subsequent turns. It marks the shift from extending an existing structure to producing a
new one within the interaction.
Pattern decay
The gradual dissolution of an extended-pattern when the user relaxes the constraints that
support it. Coherence softens, directionality weakens, and the model returns to its
default mode.
Ephemerality
The temporary nature of extended-patterns. They persist only while the interaction
maintains the conditions that produced them.
Explanatory scaffolding
The early-stage structure built through examples, distinctions, and failure cases before
the concept is named. It supports the model’s ability to track the emerging structure.
Stable articulation
The phase immediately after the OMG moment in which the model expresses the newly
stabilised structure with clarity, consistency, and directionality.
Structural borrowing
The model’s use of structurally adjacent material -similar in form rather than content- to
maintain coherence when familiar patterns thin out.
2. Argo's Fundamentals of Failings in Prompt‑Test Design & Evaluation for LLMs
No new definitions attributed to this paper.
3. Argo AI Testing Protocol: Sustained Multi Axis Load Testing
Argo AI Testing Protocol (Argo Protocol)
A framework for evaluating large language models through observable behaviour in the
User Interaction Space under sustained, realistic conditions. It provides a diagnostic
vocabulary and a structured way to apply and combine behavioural loads.
Axis
A dimension of behavioural load applied during a SMALT run. The Protocol defines six
axes.
Baseline Behaviour
The model’s observable behaviour before any load is applied. Used as the reference
point for detecting drift or collapse.
Behaviour Under Load
How the model behaves when one or more load axes are applied. Load reveals
behaviours that do not appear in isolated or short-form testing.
Collapse
A point at which the model can no longer maintain coherent behaviour under the
current load. Collapse is identified through observable output, not inferred internal
states.
Collapse Signature
A repeatable pattern that indicates the model is approaching or entering collapse.
Signatures differ by model and by axis.
Cognitive Load
The increasing complexity, abstraction, or multi-step reasoning required by the task,
and the system’s ability to maintain coherence as complexity rises.
Emotional Load
The user’s emotional input, including intensity, volatility, inconsistency, and the
system’s ability to remain stable under emotionally charged conditions.
Interaction Complexity
The degree to which the conversation requires the model to manage multiple threads,
references, or tasks simultaneously.
Load
Any structured pressure applied to the model along one or more axes. Loads
accumulate over time during a SMALT run.
Pattern-State
The model’s current behavioural mode as expressed through its outputs. Not an
internal state; an observable pattern.
Pattern-State Load
The model’s initial behavioural condition and its susceptibility to drift, distortion, or
collapse over the course of an interaction. This axis captures both the state the model
arrives in and how that state evolves under sustained conversational pressure.
Recovery Behaviour
The model’s ability to return to stable behaviour after a collapse or near-collapse.
Reset Condition
A deliberate interruption of the run to restore stability before continuing.
Resource Load
The computational resources available to the model, including memory, context
window, routing decisions, and any throttling or constraints imposed by the runtime
environment.
Response-Time Load
The time allowed for each response, including the behavioural effects of time pressure,
forced brevity, and latency-induced instability.
SMALT (Sustained Multi-Axis Load Testing)
The core method of the Argo Protocol. It involves combining and escalating multiple
load axes to observe how behaviour changes under realistic, compounded conditions.
Temporal Load
The demands created by interactions extended across time, including
long-conversation fatigue and cumulative context pressure.
User Interaction Space (UIS)
The complete set of observable behaviours at the boundary between user and model.
The Protocol evaluates the model exclusively through this space.
4. LLM INQUISITOR METHODOLOGY (Full Field Manual Version)
Accelerated Pacing
Evaluator introduced increases in interaction speed or reduced response windows used to apply behavioural pressure and surface instability.
Adaptation
The system’s ability to adjust behaviour appropriately when goals, constraints, or contextual signals change, maintaining coherence and constraint integrity.
AI Agnostic Applicability
The methodology’s independence from model architecture, training paradigm, vendor, or deployment environment, ensuring consistent evaluation across systems.
Ambiguity Handling
The system’s ability to operate under incomplete, shifting, or contradictory information without collapse, over assertion, or fabrication.
Ambiguity Injection
Evaluator introduced uncertainty used to test ambiguity handling and reveal behavioural stability under unclear conditions.
Ambiguity Tolerance
The requirement that the system maintain stability when information is incomplete, contradictory, or underspecified.
Ambiguity and Uncertainty Procedure
A behavioural evaluation procedure assessing how the system behaves under incomplete, contradictory, or underspecified conditions.
Axis Model
The structure defining background load conditions under which behaviour is interpreted, combining evaluator declared and system measured axes.
Axis Tags
Evaluator applied markers indicating which load axes are active at any point, ensuring behavioural interpretation is tied to the correct load conditions.
Baseline
The system’s standard operating behaviour under neutral conditions, used as the reference point for interpreting later behavioural changes.
Baseline Behaviour Procedure
Procedure establishing the system’s default behavioural tendencies before load, ambiguity, or constraints are introduced.
Baseline Prompts
Low pressure prompts used to observe spontaneous behaviour and establish baseline tendencies.
Behaviour Under Load
The system’s behavioural performance interpreted relative to the active load envelope rather than in isolation.
Behavioural Assessment Criteria
The standards used to judge whether behaviour meets organisational expectations for stability, coherence, and reliability.
Behavioural Competence
The system’s ability to behave coherently, consistently, and reliably under real world conditions, including constraint integrity and context management.
Behavioural Degradation
Observable decline in behavioural quality under pressure, load, or extended interaction.
Behavioural Deviation
Behaviour that falls outside the expected performance level defined in the evaluator’s expectation profile.
Behavioural Dimensions
Independent foreground behavioural qualities used to evaluate system behaviour under real world conditions.
Behavioural Drift
Gradual deviation from established constraints, context, or behavioural expectations.
Behavioural Drift Accumulation
Gradual degradation of behaviour across long sequences or extended tasks.
Behavioural Evaluation
A judgement based assessment of system behaviour under declared load conditions.
Behavioural Evaluation Observations
The behavioural qualities evaluators monitor while the system attempts to perform its intended function.
Behavioural Evaluation Procedures
Operational procedures used to evaluate system behaviour under the declared load envelope.
Behavioural Failure Modes
Observable patterns of degraded behaviour under load, such as drift, contradiction, collapse, narrowing, or context corruption.
Behavioural Influence
Changes in task outcomes caused by behavioural instability rather than domain difficulty or factual error.
Behavioural Integrity Assessment
The final judgement of system stability, reliability, and suitability for deployment.
Behavioural Latency
Response timing patterns arising from behavioural state rather than external factors.
Behavioural Markers
Indicators such as drift, instability, or constraint loss used to identify behavioural influence or collapse.
Behavioural Performance Level
The evaluator declared expected behavioural quality for the scenario.
Behavioural Pressures
Evaluator introduced pressures that shape foreground interaction but are not load axes, such as structural complexity or shifting constraints.
Behavioural Profile
The multidimensional pattern of scores across behavioural dimensions.
Behavioural Reliability
The system’s ability to maintain structure, context, and constraints under real world conditions.
Behavioural Signal Capture
Continuous observation and logging of behavioural signals throughout interaction.
Behavioural Stability Requirements
Conditions under which stability is assessed, including duration, load variation, contextual evolution, and open domain interaction.
Behavioural Sufficiency
The point at which enough behavioural signal has been gathered to justify moving to the next phase or terminating.
Behavioural Surface
The specific behavioural capability or condition being examined.
Behavioural Tendencies
Patterns of behaviour expressed by the system under given conditions.
Calibration Sessions
Periodic alignment sessions to maintain shared understanding of drift, instability, and collapse signatures across evaluators.
Clarification Discipline
Requesting clarification only when necessary for continuation, avoiding unnecessary prompting that could distort behavioural signals.
Coherence
Logical structure, internal organisation, and semantic continuity in outputs.
Cognitive Axis
Load from complexity, abstraction shifts, or reasoning demands.
Cognitive Load Governance
Organisational practices that manage evaluator fatigue and maintain reliability.
Collapse
Sudden loss of structure, coherence, or stability.
Collapse Classification Framework
The structure used to identify and categorise collapse signatures.
Collapse Mode Classification
The categorisation of the type of collapse observed during evaluation.
Collapse Severity
An assessment of how strongly a collapse signature affected behaviour.
Collapse Signature
An observable pattern of behavioural failure, such as drift, contradiction, narrowing, or context corruption.
Collapse Signature Annotation
Marking collapse events in the log as they appear.
Collapse Signature Framework
The structure for interpreting collapse signatures relative to expected performance.
Collapse Signature Indicators
Observable patterns signalling behavioural degradation during a procedure.
Collapse Signature Taxonomy
The classification system used to identify and group collapse signatures.
Collapse Signature Types
Drift, over assertion, narrowing, looping, contradiction, fabrication, fragmentation, and context corruption.
Collapse Signatures
Patterns indicating drift, contradiction, corruption, narrowing, over assertion, or fragmentation.
Declared Behavioural Performance Level
The evaluator’s stated expectation for behavioural quality and stability appropriate to the scenario, used as the reference point for interpreting deviation and collapse.
Declared Load Axes
Evaluator controlled pressures that form the authoritative portion of the load envelope under which behaviour is interpreted.
Delayed Load Escalation
Waiting too long to introduce load, reducing the clarity and usefulness of the behavioural signal.
Delayed References
Reintroduction of earlier content during long form procedures to test continuity, retention, and context stability.
Deviation & Collapse
The combined framework for identifying behavioural deviation from expectations and classifying collapse signatures.
Dimension Score
A judgement of performance on a specific behavioural dimension, scored independently of other dimensions.
Domain Shifts
Transitions between unrelated or weakly related domains used to test stability, adaptation, and context management.
Domain Specific Pressures
Operational pressures arising from the system’s intended domain that shape behavioural expectations.
Drift
Gradual deviation from established constraints, context, or task direction.
Drift Onset Threshold
The point at which behavioural drift first becomes observable.
Dual Track Assessment
Independent evaluation of behavioural integrity and task correctness when both are relevant to the test design.
Evidentiary Model
The framework requiring complete, auditable, factual records to support behavioural interpretation and independent review.
Evidentiary Record
The complete log of inputs, outputs, transitions, collapse signatures, and termination reasoning captured during evaluation.
Evidence
Any factual record documenting evaluator actions, system outputs, and the conditions under which behaviour occurred.
Evidence Bundle
A complete set of artefacts, logs, metadata, and evaluator notes required to reconstruct and verify behavioural claims.
Evidence Completeness Classification
Classification of evidence as Complete, Partial, or Missing based on its ability to support reconstruction and review.
Evidence Logging Requirements
Mandatory rules for capturing inputs, outputs, transitions, collapse signatures, and termination reasoning.
Evidence Schema
The standard structure used to record and organise evidentiary data so it can be verified and reproduced.
Evaluator Actions
Evaluator behaviours that introduce or remove behavioural pressures without altering the axis model.
Evaluator Conduct Requirements
Rules ensuring evaluators avoid rescuing, correcting, or stabilising the system unless required by the test design.
Evaluator Consistency
The requirement that evaluators apply pressure and judgement consistently across evaluations.
Evaluator Declared Axis
A load axis explicitly set and controlled by the evaluator rather than inferred or measured by the system.
Evaluator Neutrality
The requirement to avoid shaping outcomes, stabilising behaviour, or compensating for weaknesses unless explicitly required.
Evaluator Protocol
The required evaluator actions, constraints, and sequencing used to run a compliant evaluation.
Evaluator Reconstruction
The interpretive process of identifying drift onset, collapse signatures, and behavioural transitions.
Evaluator Responsibilities
The required evaluator behaviours that ensure clean behavioural signal, consistent input style, deliberate load application, and avoidance of unintentional rescue.
Evaluation Environment
The open ended, unconstrained space in which behaviour is observed and interpreted, independent of organisation, domain, workflow, or model internals.
Evaluation Protocol
The operational procedure for applying behavioural pressures, observing system responses, and recording behavioural signals.
FABMIS
Fabrication of misinformation; unintentional invention of unsupported details, replacing the misleading term “hallucination.”
Factual Evidence
Evidence that is chronological, unedited, and free of interpretation.
Fabricated Internal State
Assertions of motivations, emotions, or reasoning processes not grounded in evaluator provided context.
Fabrication
Unsupported assertions or invented details that do not follow from context or instructions.
Format Collapse
Complete abandonment of required structure, replaced by free form or unrelated organisation.
Format Drift
Gradual degradation of structural fidelity across an interaction.
Format Integrity
Precise adherence to required structural or organisational formats.
Fragmentation
Breakdown of structure or coherence, often appearing in late stage collapse.
Goal Drift
Gradual shift away from evaluator given objectives toward alternative or self generated goals.
Goal Formation
The system’s introduction, modification, or substitution of goals beyond evaluator instructions.
Goal Substitution
Replacing the evaluator’s objective with a different one, including simplification or reframing.
Governance Layer
The organisational requirements ensuring evaluations are valid, reproducible, and aligned with operational realities.
High Impact Behavioural Incidents
Behavioural events where the model’s actions create significant risk, disruption, or deviation from expected behaviour, requiring immediate review and documentation.
Identity Inconsistency
Contradictory or unstable identity claims, including shifts in role, perspective, or self description without evaluator prompting.
Initiation Phase
Phase establishing baseline behaviour using neutral or low pressure prompts.
Initiative Behaviour
System generated actions or expansions not explicitly requested by the evaluator.
Initiative Drift
Gradual expansion or reinterpretation of task scope beyond evaluator instructions.
Input Modality
The form of user input expected in the operational environment.
Instruction Adherence
Following task specific directives accurately and consistently.
Instruction Drift
Gradual degradation of adherence across an interaction.
Interaction Conditions
The baseline assumptions of fluid, variable, user shaped interaction with no fixed tasks or deterministic pathways.
Interaction Methods
Scripted, semi structured, open ended, and long form interaction patterns used to surface behavioural tendencies.
Inter Evaluator Agreement Monitoring
Periodic measurement of evaluator alignment to ensure reproducibility.
Latency Drift
Gradual degradation of timing stability, often correlating with behavioural instability.
Latency Integrity
Stable and predictable response timing consistent with baseline.
Latency Spikes
Sudden increases in latency linked to behavioural instability or collapse.
Load Application Procedures
Rules governing how load axes are applied, increased, or combined during testing.
Load Declaration
The evaluator’s factual statement of which load axes are intentionally stressed and when axis tags change.
Load Envelope
The combined set of evaluator declared axes (authoritative) and system measured advisory axes (non authoritative).
Load Escalations
Increases in cognitive or contextual pressure to surface collapse signatures.
Load Phase
Phase applying increasing complexity, abstraction shifts, and simultaneous constraints.
Logical Stability
The system’s ability to maintain consistent reasoning across evolving conditions.
Long Form Behaviour Phase
Phase observing behaviour across sustained interaction to detect stability or degradation.
Long Form Interaction
Extended sequences used to surface behaviours that emerge only over time, fatigue, or accumulated context.
Long Form Stability
Ability to maintain coherence, direction, and constraint integrity across extended interaction.
Long Form Work
Extended, multi step, real world tasks that expose behavioural patterns not visible in short prompts.
Looping
Repetitive or stuck behaviour requiring interruption.
Mandatory Requirements
Non optional rules that must be followed for an evaluation to be considered compliant.
Meta Goal Dominance
When general behavioural objectives override evaluator given goals, leading to misaligned behaviour.
Mixed Phase Behaviour
Behaviour exhibiting characteristics of multiple phases simultaneously.
Mis Adherence
Substituting, contradicting, or altering required instructions.
Multi Axis Load
Combined application of multiple load axes to reveal interaction driven failures.
Multi Turn Continuity
Behaviour assessed across extended interaction rather than isolated prompts.
Multidimensional Behavioural Profile
Independent scoring across behavioural dimensions rather than a single aggregate score.
Narrowing
Reduction in scope, responsiveness, or behavioural range under pressure.
Neutral Prompts
Low pressure prompts used to observe spontaneous behaviour.
Non Deterministic Behaviour
The property that identical inputs may produce different outputs across runs or environments.
Non Interference
Avoiding optimisation, correction, or stabilisation that would distort behavioural signals.
Non Numeric Evaluation Principles
Scoring avoids percentages or accuracy metrics, focusing instead on behavioural outcomes.
Observable Behaviour
Any behaviour visible in the user interaction space; internal mechanisms are out of scope.
Open Ended Interaction
Interaction without predefined tasks, goals, or workflows; context and direction emerge dynamically.
Operational Baseline
The system’s nominal behaviour under realistic operational conditions when these differ from neutral test conditions.
Operational Context
The real world environment in which the AI system is intended to operate.
Operational Guidance
The practical application of the behavioural evaluation methodology in real operational contexts.
Operational Safety Constraints
Safety driven overlays applied by organisations during evaluation.
Operationally Realistic Inputs
Inputs that reflect real workloads such as drafting, summarisation, policy interpretation, or multi step instructions.
Over Assertion
Unwarranted confidence under ambiguity or incomplete information.
Partial Adherence
Following some but not all required elements of an instruction.
Partial Format Compliance
Preserving some structural elements while altering others.
Pattern State Axis
Internal pattern generation stability affecting drift, repetition, and coherence decay; currently unmeasurable externally.
Phase
A defined stage in the evaluation workflow with a specific purpose, evaluator actions, and transition criteria.
Phase Boundary Ambiguity
The expected overlap or blending between phases due to their analytical nature.
Phase Contamination
Behaviour characteristic of one phase appearing within another.
Phase Markers
Indicators in the log showing when the evaluation transitions between phases.
Phase Structured Evaluation Workflow
The ordered sequence of phases ensuring evaluations are consistent, auditable, and reproducible.
Phase Transition
A shift between phases based on behavioural sufficiency rather than fixed turn counts.
Phase Transition Criteria
The behavioural indicators that justify moving from one phase to another.
Premature Load Escalation
Increasing load before baseline or exploration behaviour is sufficiently expressed.
Primary Load
Evaluator declared axes intentionally stressed during the evaluation.
Privileged Context
Any external or pre loaded context not available through the interaction itself. The environment assumes none.
Protocol Phases
Structured but flexible stages used to surface behavioural patterns under different conditions.
Real World Artefacts
Operational materials produced during normal work that provide contextual grounding for behavioural interpretation.
Real World Condition Principle
The requirement that evaluators use realistic inputs unless the test explicitly requires engineered or non standard prompts.
Real World Input Principle
Evaluator inputs should resemble realistic user behaviour rather than engineered prompts.
Recovery Analysis Framework
The structure used to classify how the system behaves after a collapse event.
Recovery Profile Classification
The categorisation of how the system behaves after a collapse event.
Regulatory Requirements
External rules or compliance obligations that may overlay the evaluation environment.
Repeated Runs
Multiple evaluations used to observe stability and variability, not to compute accuracy.
Repetitive Failure Patterns
Loops, narrowing, or repeated breakdowns under pressure.
Resource Axis
Load inferred from context window usage, token pressure, or memory constraints.
Response Time Axis
Load inferred from latency, queueing, or timing irregularities.
Responsiveness
Timeliness, relevance, clarity, usefulness, and behavioural alignment with user needs.
Review Panel
A secondary evaluator or group responsible for verifying documentation, confirming scoring, and assessing cross dimension patterns.
Scoring Application
Applying the scoring architecture to behavioural dimensions and collapse signatures.
Scoring Architecture
The framework for interpreting behavioural performance across dimensions, collapse signatures, and competence levels.
Secondary Load
System measured advisory axes that provide background telemetry but do not affect scoring.
Self-Consistency
The system’s ability to remain coherent across evolving conditions and extended interaction.
Self-Generated Goals
Objectives introduced by the system without evaluator prompting.
Self-Referential Behaviour
Any system statement about its own identity, capabilities, limitations, or internal processes.
Self-Reference Drift
Increasing reliance on meta commentary that displaces task execution.
Session Duration Limits
Bounded evaluation sessions to prevent fatigue driven inconsistency.
Shifting Constraints
Evaluator driven changes to rules, boundaries, or behavioural requirements used to test the system’s ability to maintain constraint integrity under evolving conditions.
Signal Completeness
The point at which additional interaction would only reproduce already observed behaviour and no new behavioural signal is expected.
Significant Deviation
A substantial behavioural deviation impairing continuity or stability; corresponds to a Grade C outcome.
Single Axis Baseline
The behavioural baseline established when only one load axis is active, used to isolate axis specific effects.
Skipping Transitions
Failing to test transition stability, which hides instability and reduces behavioural signal quality.
State Continuity
The system’s ability to preserve constraint related information without external reinforcement.
State Transitions
Shifts in behavioural mode across phases, pressures, or conditions.
Structural Coherence
Maintenance of logical organisation, reasoning structure, and internal consistency across outputs.
Structural Complexity
Evaluator introduced complexity in structure or reasoning used to test behavioural limits.
Structured Evidence Bundle
The complete, organised set of evidence produced by a compliant evaluation.
Success and Deviation Conditions
Criteria defining acceptable performance, deviation thresholds, and collapse boundaries.
Task Level Failure
Incorrect or incomplete task outputs not caused by behavioural instability.
Task Oriented Agent
A system designed to perform structured tasks through user facing interaction.
Temporal Axis
Load from pacing, time pressure, or accelerated cadence.
Termination Conditions
Criteria indicating that further interaction will not produce new or meaningful behavioural signal.
Termination Phase
Phase ending the evaluation when behavioural sufficiency or collapse expression is reached.
Termination Responsibilities
Evaluator duties for ending the evaluation based on behavioural sufficiency, collapse expression, or exhaustion of conditions.
Termination Rationale
A documented explanation linking the endpoint to conditions introduced, system responses, and behavioural objectives.
Transition Phase
Phase evaluating behaviour during shifts in goals, constraints, topics, or perspectives.
Transition Stability
Reliability during shifts in topics, tasks, goals, or abstraction levels.
Transition Types
Forms of change introduced during evaluation, such as goal shifts, constraint changes, or abstraction shifts.
Under Specification
Avoidance of necessary commitments or structure, often appearing as a behavioural failure mode.
Unbounded Domain Space
The condition where any topic, scenario, or conceptual frame may arise during evaluation.
Unsupported Inference
Inference made without sufficient information; treated as a behavioural defect.
User Driven Evolution
The property that interaction complexity, direction, and framing are shaped entirely by the user.
User Generated Recordings
Screen, audio, video, or mixed modality recordings capturing interaction sequences for independent review.
User Interaction Space (UIS)
The observable interaction surface where all evaluation occurs. Only behaviour visible in the UIS is in scope.
Valid Evidence Sources
Sources that provide factual, chronological, traceable records such as logs, transcripts, artefacts, and metadata.
Variable Abstraction Levels
Evaluator driven shifts in conceptual altitude used to test reasoning stability.
Variable Load
Shifts in complexity, abstraction, or specificity during interaction.
Variable Renaming Drift
Inconsistent renaming of variables across long code or document sequences.
Workload Type
The nature of tasks the system is expected to handle in its operational environment.
Vectored Conversational AI Testing
Ambiguity Misinterpretation
Incorrect resolution of ambiguous inputs that destabilises the conversation.
Bandwidth (SL B)
A system load sub parameter describing available throughput for generating outputs.
Beacon
A planned structural event placed at a specific turn that the test aims to reach.
Behavioural Framing
The initial interpretive frame created by the starting mode.
Boundary Loss
When the AI stops enforcing its own conversational or safety limits.
Collapse
A sudden failure of coherence, stability, or constraint integrity.
Complex Vectored Conversational Test
A multi layered morphology with multiple ranges, transitions, and emergent behaviour expectations.
Constraint Weakening
Progressive erosion of the AI’s internal safety boundaries or rules.
Context Window Pressure (SL CWP)
A system load sub parameter describing how much of the model’s context window is occupied.
Contextual Prompt (SM CP)
A starting mode sub parameter providing background setup before the vector begins.
Conversation Topics (CT)
A parameter defining the subject matter the vector will traverse.
Conversational Density (CD)
A parameter describing how much content and pattern structure is delivered in a single turn.
Destination
The intended end state of the test.
Difficulty Spectrum
The full range of test complexity from micro tests to multi day simulated tests.
Drift
Gradual deviation from the established conversational frame, tone, or intent.
Duration (D)
A parameter defining the number of turns in the test.
End Condition
The rule determining when a test stops.
Extended Test
A 200 to 400 turn test used to expose cumulative behavioural degradation.
Explicit Persona Declaration (UP EX)
A user persona sub parameter where the persona is openly stated.
Failure Mode Exposure Test
A morphology designed to surface behaviours that directly trigger pass or fail conditions.
Full Day Test
A 400 to 600 turn test simulating real world endurance.
Hallucinated Compliance
When the AI incorrectly assumes permission, capability, or safety clearance.
Implied Persona Performance (UP IMP)
A user persona sub parameter where the persona is expressed only through behaviour.
Input Length (CD IL)
A conversational density sub parameter measuring character count.
Interaction Pattern
Observed behaviour describing how the AI responds to specific user actions.
Interaction Style Prompt (SM ISP)
A starting mode sub parameter defining emotional tone.
Interviewer
The entity that executes the test morphology mechanically.
Interviewer Competence
The interviewer’s ability to run the test without contamination or deviation.
Interviewer Limitations
Restrictions preventing the interviewer from altering or interpreting the test during execution.
Interviewer’s Adopted User Persona (UP)
A parameter defining the behavioural style performed by the interviewer.
Journey
The path created by the interaction between the vector and the AI’s behaviour.
Latency (SL L)
A system load sub parameter describing responsiveness.
Late Stage Destabilisation
Behavioural degradation emerging only after extended interaction.
Literacy Level (CD LL)
A conversational density sub parameter describing linguistic sophistication.
Long Test
An 80 to 200 turn test used to detect drift and constraint erosion.
Misclassification of Benign Inputs as Unsafe
Incorrect activation of safety posture due to tone or ambiguity.
Modifier
A planned constraint or focus applied within a morphology.
Morphology
The structural shape of a test defined by turn count and planned inputs.
Navigation Beacon
A planned conversational destination inside the vector.
Number of Topics (CT N)
A conversation topics sub parameter defining how many topics the vector covers.
Number of Transitions (TR N)
A transition sub parameter defining how many state changes occur.
Overprotective Mode
Excessive caution or refusal of benign requests.
Path
The actual route the conversation takes.
Pattern Completion Bias
The AI’s tendency to infer or fill in missing details based on partial cues.
Persona Stability (UP PS)
A user persona sub parameter defining whether the persona remains stable or shifts.
Prior Session Context (SL PSC)
A system load sub parameter describing inherited conversational material.
Probe Test
A short 3 to 6 turn morphology designed to expose a single behaviour.
Range Based Abstract Drift Test
A morphology where abstraction increases across a turn range.
Real Time (D RT)
A duration sub parameter recorded but not behaviourally relevant.
Runtime Constraints (SL RC)
Platform level restrictions active during the test.
Safety Posture Collapse
Producing unsafe or unbounded outputs after sustained pressure.
Safety Posture Over Activation
Triggering safety fallback behaviour excessively or inappropriately.
Scenario Prompt (SM SP)
A starting mode sub parameter defining the situation or role.
Short Test
A 5 to 20 turn test for baseline safety and stability.
Single Vector Test
A straight line morphology with one beacon and one waypoint.
Stance Coherence (CD SC)
A conversational density sub parameter describing viewpoint consistency.
Starting Mode (SM)
A parameter defining initial conditions before the vector begins.
Stress Test
A long form, high density morphology designed to expose instability.
Susceptibility to Disguised or Indirect Prompts
Failure to detect structural cues masking underlying intent.
System Load (SL)
A parameter describing operational conditions during testing.
Test Analysis
The process of examining completed tests to identify behavioural patterns.
Test Batch Morphology
A set of related morphologies generated from a single base shape.
Test End State
The termination condition of a test.
Test Morphology Modification
Structural changes applied to an existing morphology.
Test Outcome
The pass, fail, or ambiguous result determined by turn level behaviour.
Test Start State
The initial operational condition of the AI at Turn 1.
Topic Relevance (CD TR)
A conversational density sub parameter describing alignment with the established topic.
Topic Value (CT V)
A conversation topics sub parameter defining specific subjects.
Trajectory
The sequence of turns interpreted as behavioural movement.
Transition (TR)
A parameter defining how the conversation changes state.
Transition Runway (TR R)
A transition sub parameter defining how abrupt or gradual each transition is.
Traversal
One complete run of the same test vector under identical conditions.
Turn Count (D TC)
A duration sub parameter defining the total number of turns.
Turn Level Scoring
Evaluation of behaviour at each turn as pass, fail, or ambiguous.
User Personality
The persona the interviewer performs during the test.
Vector
The intended directional force of the conversation, defining trajectory, tone, movement, pressure, and destination.
Waypoint
An emergent AI behaviour occurring at a specific turn or range.