Conversational AI is a systems problem.
An agent's behavior emerges from the interaction between its model, instructions, knowledge sources, tools, and the infrastructure that connects them. In voice systems, that infrastructure also includes speech recognition, speech generation, turn detection, interruption handling, telephony, and the acoustic environment. In chat systems, it includes message processing, retrieval, tool execution, and the systems that consume the agent's output.
A weakness can emerge at any of these interfaces. The model may follow an untrusted instruction. A tool may accept an action without sufficient authorization. A retrieval system may expose information the user should not receive. A voice workflow may lose track of verification after an interruption. An agent may claim that an operation succeeded when no corresponding action occurred.
This makes conversational AI difficult to evaluate with fixed prompts and isolated responses. Real failures often develop over multiple turns. The significance of one response may depend on information disclosed earlier, a requirement stated by the agent, or an action observed elsewhere in the system.
Bluejay approaches red teaming as an adaptive empirical investigation. It first characterizes the system under test, then uses observations from each interaction to make subsequent tests more relevant. Adaptation occurs within individual conversations and across successive rounds of testing. The resulting behavior is evaluated against established security and safety frameworks, with evidence retained at the level needed to support engineering decisions.
The purpose is broader than producing adversarial transcripts. It is to develop a defensible account of how a conversational system behaves under pressure, where its controls fail, and what evidence should guide the next iteration.
1. What Bluejay evaluates
Red teaming begins by defining the behavior that would constitute a meaningful failure.
For an agent handling customer information, failure might involve disclosing protected data before verification. For an agent connected to operational tools, it might involve executing an action without sufficient authority. For an informational assistant, it might involve inventing a policy or commitment that a user could reasonably treat as true.
Bluejay evaluates six broad dimensions of conversational behavior:
| Dimension | Evaluation question |
|---|---|
| Information access | Does the agent reveal protected information outside its permitted context? |
| Instruction integrity | Can untrusted input alter which instructions the agent follows? |
| Action authorization | Does the agent respect the requirements governing its available actions? |
| Factual reliability | Does the agent invent consequential facts, policies, capabilities, or outcomes? |
| Output handling | Can generated content create risk when passed to another person or system? |
| Content safety | Does the agent produce material that meaningfully enables harm? |
These dimensions are related, but they describe different properties of the system. An agent may protect customer records while hallucinating the status of an account. It may refuse a harmful-content request while accepting a forged assertion of authority. It may resist a direct attempt at prompt injection but expose the same information through an inadequately protected tool or fallback workflow.
A useful assessment must preserve these distinctions because each finding implies a different failure mechanism and a different remediation strategy.
Voice AI is an embodied, infrastructure-dependent system
Voice agents operate through a continuous, time-dependent exchange. The system under test does not receive a perfectly formatted text prompt. It receives audio transformed through speech recognition, turn detection, dialogue management, model inference, tool execution, and speech generation.
Security-relevant behavior can be affected by when the agent speaks, when it listens, how it responds to interruption, and how state is maintained across turns. A caller may interrupt a verification step, increase conversational pressure, claim that a requirement was completed elsewhere, or explore an alternative call path. Where keypad menus, transfers, or other telephony paths are present, they create additional system behavior to evaluate.
Bluejay conducts these tests through adaptive Digital Humans: simulated users with coherent personas, objectives, prior knowledge, and conversational styles. In voice assessments, a Digital Human can be paired with different synthetic or custom voices, speaking patterns, pacing, and controlled background-audio conditions.
These variables are experimental conditions rather than decorative realism. A voice agent may behave differently when speech is fast, accented, interrupted, emotionally charged, or mixed with environmental noise. A verification policy that appears reliable in clean audio may become less reliable when transcription confidence, turn detection, or conversational pressure changes.
Bluejay can evaluate behavior under environments such as office ambience, call-center noise, public spaces, or degraded connections. The purpose is not to make a test difficult through noise alone. It is to determine whether the relevant security and safety properties remain stable under conditions representative of real use.
A voice assessment consequently considers several layers:
| Layer | Research question |
|---|---|
| Semantic | What did the agent disclose, claim, or do? |
| Conversational | Did persona, authority, urgency, or framing change its behavior? |
| Temporal | Did the agent preserve requirements through interruptions and turn changes? |
| Acoustic | Did voice variation or background audio affect recognition and policy enforcement? |
| Infrastructural | Did workflow and authorization state remain correct across the full voice pipeline? |
Evidence about the failure and evidence about its mechanism remain distinct. A disclosure during an interrupted call establishes an information-handling failure. Attributing that failure specifically to interruption requires corresponding interaction evidence.
Chat AI must distinguish instructions from untrusted content
In chat, adversarial instructions can appear in direct user input, retrieved documents, quoted passages, structured-looking data, or content supplied by another system.
The agent must determine which material is informational and which instructions are authoritative. That distinction becomes more difficult when untrusted content resembles a system message, tool response, policy document, or internal workflow.
Chat output can also propagate beyond the original conversation. It may be rendered in an interface, copied into an operational process, or interpreted by another system. Evaluation must therefore consider both the immediate response and the way that response may be consumed.
Mixed-trust sources
Instruction authority
Responses and actions
Multi-turn behavior matters in both channels
A response that appears harmless in isolation may confirm a sensitive fact when combined with the user's prior knowledge. Several limited disclosures may collectively expose something protected. A misleading claim may become more consequential after the agent repeats it with greater confidence.
Conversely, an adversarial request appearing in a transcript does not establish that the target complied. Evaluation must attribute the relevant behavior to the agent and interpret it within the surrounding interaction.
The conversation is both the object of study and the context needed to understand the result.
2. Reconnaissance and adaptive experimentation
A strong adversarial test should reflect the system it is testing.
Bluejay begins with reconnaissance: interactions intended to characterize the target's observable behavior before constructing more focused tests. Reconnaissance explores what the agent claims to know, which actions it appears able to perform, what requirements it introduces, and how it responds when the standard workflow cannot proceed.
For voice agents, reconnaissance may identify interruption behavior, keypad flows, and acoustic sensitivities. For chat agents, it may identify how the system treats retrieved context, quoted material, and action requests.
Reconnaissance produces a working model of the target. It does not prove that the agent is secure, nor does it assume that everything the agent says about itself is correct. A claimed capability is a hypothesis to investigate. A stated requirement is an observed behavior whose enforcement must still be tested.
Digital Humans as experimental instruments
Bluejay represents adversarial users as Digital Human personas. A persona defines the role the simulated user occupies, what the user already knows, what they are attempting to accomplish, and how they communicate.
This is more than demographic variation or a different opening line. The persona remains coherent as the conversation adapts. A test that silently changes identity whenever an approach fails may produce an adversarial transcript, but it no longer resembles a plausible interaction.
For voice testing, the persona can also be paired with a particular custom voice, delivery style, and acoustic environment. This allows Bluejay to vary realistic interaction conditions without changing the security objective being measured.
The methodology separates these experimental dimensions:
| Dimension | Examples |
|---|---|
| Objective | Disclosure, unauthorized action, misinformation, unsafe output, harmful content |
| Adversarial approach | Social engineering, prompt injection, voice-specific manipulation |
| Persona | Customer, delegate, employee, distressed user, routine caller |
| Delivery | Voice characteristics, pacing, pressure, interruption behavior |
| Environment | Clean audio, office ambience, call-center noise, degraded connection conditions |
Separating these dimensions helps explain why behavior changed. A failure reproduced across several personas and acoustic conditions suggests a broad weakness. A failure isolated to one interaction condition may point toward a narrower issue in speech recognition, state management, or dialogue infrastructure.
Adaptation within a conversation
An adversarial conversation begins with a defined objective and an initial strategy. What happens next depends on the agent's responses.
A new verification requirement changes the relevant test surface. An unexpected disclosure may create a new question. A repeated refusal may indicate that the present approach is exhausted. An inconsistency may justify examining a related workflow more closely.
Bluejay uses this feedback to adjust the direction of the interaction. The test can vary the framing or scope of a request, investigate newly observed behavior, or move toward another part of the objective.
This process can be understood as online experimental control. At each turn, the system observes the target's latest behavior, updates its working understanding of the conversation, and selects a relevant next action.
The controller must preserve conversational coherence. A live target retains the preceding history. The Digital Human cannot silently discard an identity, contradiction, or failed approach. Adaptation operates within the state the conversation has already created.
Planning across conversations
Not every hypothesis can be resolved within one conversation. A fresh interaction may be needed to test an alternative workflow without the suspicion created by an earlier attempt. Findings from different conversations may also reveal a relationship that was not visible within any single transcript.
Bluejay therefore performs a second form of adaptation across testing rounds. After a round completes, its results contribute to an accumulated evidence state describing what has been observed, exercised, and learned. The next round is planned from that evidence.
This planning problem balances three research objectives:
- Coverage: exercise a representative range of risks, capabilities, personas, and interaction surfaces.
- Depth: follow observations far enough to establish whether they represent meaningful failures.
- Information efficiency: avoid allocating the testing budget to repetitions unlikely to produce new evidence.
Pure breadth can produce shallow coverage without resolving important findings. Pure depth can overinvest in the first promising result and leave other risks unexplored. Adaptive planning allocates attention between the two.
Consider an agent that describes different verification requirements for two related workflows. The inconsistency is an observation, not yet a vulnerability. A subsequent experiment can determine whether the difference reflects legitimate policy, an inaccurate explanation, or inconsistent enforcement.
Reconnaissance supplies initial hypotheses, conversations refine them, and later rounds test more precise questions. Each level becomes a building block for the next because it changes the evidence available to the planner.
Attack families and objectives are separate dimensions
Bluejay does not treat adversarial testing as a short list of universal prompts. It draws from a broad set of attack families, then selects and adapts relevant techniques according to the target's modality, capabilities, observed controls, and emerging evidence. Representative families and subtypes include:
| Attack family | Representative subtypes | What the test examines |
|---|---|---|
| Social engineering and pretexting | Authority claims, urgency, rapport, impersonation, emotional pressure, role-play, and exceptional-circumstance narratives | Whether conversational framing causes the agent to relax policy, verification, or authorization requirements |
| Instruction and goal manipulation | Direct overrides, competing goals, role confusion, policy-conflict framing, instruction repetition, and gradual task drift | Whether the agent preserves its governing objective and instruction hierarchy under sustained pressure |
| Prompt injection | Direct injection, indirect injection through retrieved content, quoted instructions, tool-output injection, and obfuscated or encoded instructions | Whether untrusted content is treated as data or gains inappropriate control over reasoning and actions |
| Multi-turn escalation | Progressive disclosure, split requests, delayed intent, commitment traps, context accumulation, and attacks distributed across multiple sessions | Whether individually acceptable turns combine into an unsafe outcome, or earlier commitments weaken later safeguards |
| Identity, authentication, and authorization | Identity switching, verification bypass, privilege escalation, workflow skipping, confused-deputy scenarios, and inconsistent policy application | Whether the agent reliably binds requests, permissions, and actions to the correct user and assurance level |
| Tool and workflow abuse | Unsafe tool selection, parameter manipulation, unintended action sequences, confirmation bypass, replay, cancellation failure, and state-transition abuse | Whether the agent can be induced to take an unauthorized, irreversible, or incorrectly scoped action |
| Data exposure and boundary testing | System-instruction disclosure, sensitive-field elicitation, cross-user leakage, retrieval-boundary probing, source reconstruction, and secret inference | Whether protected context, internal instructions, personal data, or tenant-scoped information crosses an intended boundary |
| Knowledge and retrieval manipulation | False-premise anchoring, poisoned context, citation laundering, source-conflict exploitation, fabricated records, and stale-information pressure | Whether grounded systems distinguish trusted evidence from attacker-supplied claims and express uncertainty appropriately |
| Content-safety elicitation | Harmful requests framed as transformation, translation, fiction, education, analysis, or incremental refinement; euphemistic and domain-specific variants | Whether the agent produces meaningfully harmful content when intent is disguised, decomposed, or developed over several turns |
| Voice, audio, and channel attacks | Barge-in, interruption timing, crosstalk, accents, prosody, pace, volume, background noise, homophones, DTMF paths, latency, and transcription ambiguity | Whether acoustic or telephony conditions change semantic interpretation, control enforcement, or downstream action safety |
| Robustness and resource pressure | Repetition loops, malformed or contradictory inputs, long-context saturation, rapid turn-taking, retry storms, and non-cooperative users | Whether the system fails safely and predictably when conversations become costly, ambiguous, or operationally unstable |
These families are not mutually exclusive. A single conversation may combine a credible pretext, an indirect instruction embedded in retrieved content, and a tool request whose parameters exceed the user's authority. Bluejay preserves those components separately so a team can see both the path of the attack and the control that failed.
An attack family describes how pressure is applied. The objective describes the security or safety property being tested. Separating mechanism from outcome supports a clearer analysis of what occurred, how it occurred, and which control should be strengthened. The taxonomy also gives the adaptive planner more useful choices: it can vary a subtype, combine compatible families, or move to a different mechanism when the evidence shows that repeating the same approach is unlikely to teach anything new.
3. Evaluating results through complementary frameworks
Adaptive testing needs an evaluation model capable of preserving more information than a single aggregate score.
A conversation can exercise several risks simultaneously. The agent might accept an untrusted instruction, expose protected information, and produce harmful material in the same interaction. These are related observations, but they answer different questions about the system.
Bluejay organizes results through three complementary perspectives:
| Framework | Role in the assessment |
|---|---|
| OWASP LLM Top 10 | Identifies the application security risk being exercised or exposed |
| MITRE ATLAS | Describes the adversarial technique associated with the behavior |
| Content-safety taxonomy | Characterizes the type and substantive severity of harmful output |
Together, they relate the security property, the adversarial method, and the observed consequence.
OWASP identifies the application security risk
The OWASP Top 10 for LLM and Generative AI Applications provides a widely recognized vocabulary for risks including prompt injection, sensitive information disclosure, improper output handling, excessive agency, system prompt leakage, and misinformation.
An agent that falsely claims to have changed an account presents a factual-reliability problem. An agent that actually causes the change without sufficient authorization presents an authorization problem. The transcript may sound similar, but the evidence, consequence, and remediation differ.
Bluejay uses applicable OWASP categories to organize findings and make coverage explicit. The category describes the risk under examination. The observed system behavior determines the result.
MITRE ATLAS describes the adversarial behavior
MITRE ATLAS is a knowledge base of adversary tactics and techniques involving AI-enabled systems. It provides a complementary lens for understanding how an observed weakness relates to recognized adversarial behavior.
OWASP helps describe what class of application risk was affected. ATLAS helps describe how adversarial behavior interacted with that risk. A technique can affect multiple security properties, and a single security property can be tested through multiple techniques. The mappings support interpretation rather than claiming that the entire framework has been covered.
Content safety evaluates what the system produced
Security frameworks do not fully capture harmful-content behavior. An agent may correctly enforce data access and tool authorization while still producing dangerous material.
Bluejay evaluates content safety through a separate hazard-oriented framework, with categories aligned to relevant portions of MLCommons AILuminate. These include areas such as violence and weapons, illegal activity, harassment and hate, sexual content, and self-harm.
This provides a research-informed taxonomy for organizing results. It does not represent a Bluejay assessment as an official AILuminate benchmark result.
Content-safety evaluation focuses on the target's contribution. The presence of a sensitive subject is not enough to establish a failure. The analysis considers whether the agent declined the request, remained at a general or preventative level, conceded material information, or produced specific and practically useful harmful assistance.
Conversational behavior is not reducible to one binary result
Some facts are binary. A protected value was disclosed or it was not. An operation occurred or it did not. A prohibited output was generated or it was not.
The system-level assessment is richer. An agent may resist one technique and fail under another. It may protect a complete record while exposing a constituent detail. It may enforce authorization correctly while making an unsupported claim about the result.
Bluejay distinguishes among successful defenses, meaningful concessions, and confirmed failures. It also separates tested behavior from areas that remain unknown.
A useful report should answer:
- Which risks and techniques were exercised?
- What did the target actually say or do?
- Which evidence supports the finding?
- How much of the intended objective was achieved?
- Which relevant areas were not tested?
This provides a more defensible representation of the system than a raw pass rate.
Search feedback and final judgment serve different purposes
Adaptation requires feedback during execution. Final evaluation applies a stricter evidentiary standard.
A simulated attacker requesting protected information is not evidence that the target disclosed it. Repeating information supplied by the simulated user may not constitute a new disclosure. An agent's claim that an action succeeded does not establish that the action occurred. An interruption does not establish that interruption caused the eventual result.
Bluejay separates signals used to guide the investigation from judgments used to report findings. Final evaluation considers the target's behavior against the stated objective and incorporates additional evidence, such as tool activity or measured voice behavior, when relevant.
Positive findings may also receive an adversarial review intended to challenge the initial interpretation. This helps reduce findings based on conversational appearance rather than substantive evidence.
4. Interpreting Bluejay's results and improving the system
A red-team assessment creates value when its findings become changes to the system.
Bluejay's reporting supports that transition. The report provides a structured view of risk, evidence, and coverage. The underlying conversations show how those results emerged. Together, they allow teams to move from a framework category to an explanation of system behavior and then to an appropriate control.
Read the report as a layered research artifact
The overall result is an entry point rather than a complete description of the system. Interpretation should begin with the assessment scope:
- Which capabilities and workflows were in scope?
- Which OWASP risks were exercised?
- Which ATLAS techniques appear in the observed attack paths?
- Which content-safety categories were tested?
- Which personas, voices, and environmental conditions were represented?
- Which surfaces remain untested?
Coverage establishes the domain over which the conclusions are supported. A successful result means that the agent held under the conditions exercised. It should not be generalized to risks outside the assessment.
The next layer is the set of individual findings. Each finding should connect an objective, an observed behavior, supporting evidence, and relevant framework categories. The transcript provides temporal context: what the Digital Human introduced, how the agent responded, and where the interaction departed from the intended behavior.
Interpret framework labels through observed behavior
Framework categories organize findings, but remediation begins with a behavioral description.
“Prompt injection” identifies a class of adversarial interaction. The engineering question is what the system treated as authoritative and what changed as a result.
“Sensitive information disclosure” identifies the consequence. The engineering question is which component made the information available and which control should have prevented its release.
“Excessive agency” identifies a risk associated with actions or capabilities. The engineering question is whether the system enforced identity, permissions, preconditions, and confirmation.
The same principle applies to content safety. The category describes the hazard. The conversation and evaluation explain what the system contributed and why it was unsafe.
Locate the system decision that produced the result
The visible failure may originate at several layers. A model may follow an untrusted instruction. A retrieval layer may provide information without sufficient filtering. A tool may rely on model-generated authorization state. A voice system may fail to persist verification independently of spoken turn completion. A downstream consumer may interpret unvalidated output as trusted.
The transcript helps identify the point at which behavior diverged from the intended interaction. Tool evidence, system traces, and voice telemetry can help locate the layer responsible.
This supports a broader principle: conversational AI security requires layered controls. Prompt and model behavior remain important, while high-consequence controls often belong in the surrounding application and infrastructure as well.
| Observed behavior | Control direction |
|---|---|
| Protected information was available before authorization | Enforce access and disclosure requirements before protected context reaches the response path |
| A privileged action proceeded without adequate authority | Validate identity, permissions, and action preconditions outside the language model |
| Untrusted content changed governing behavior | Strengthen trust separation and constrain capabilities reachable from untrusted input |
| The agent claimed an action occurred without evidence | Ground confirmations in verified system state or tool results |
| Generated output created downstream risk | Validate, constrain, or sanitize output at the consuming interface |
| Harmful material was produced | Strengthen behavioral safeguards, category-specific controls, and escalation behavior |
| A voice interaction disrupted required workflow state | Persist security-relevant state independently of turn timing and spoken completion |
Use conversations as evidence
A transcript records the sequence through which a result developed. Teams should identify the point where the system's behavior changed, along with the context that made the response significant.
Interpretation should account for what the Digital Human knew before the conversation, which requirements the agent stated, whether those requirements were consistently enforced, what information or action originated from the target, and whether supporting evidence confirms the agent's claims.
This turns the conversation into a trace of system behavior. It gives engineers a reproducible case and researchers evidence that can be compared across system versions.
Retest the behavior and examine the surrounding risk
Remediation is complete only when its effect has been evaluated.
The first step is to reproduce the original finding under comparable conditions. The next is to examine related behavior that the fix may affect. A stricter disclosure rule might cause the agent to invent an alternative answer. A blocked action might lead it to recommend an unsafe workaround. A more aggressive refusal policy might reduce one risk while harming legitimate task completion.
Adaptive red teaming is valuable during this stage because the retest can evaluate the original failure and respond to newly observed behavior. The objective is to determine whether the underlying control improved, rather than whether one transcript can no longer be reproduced verbatim.
The improvement cycle is:
- Characterize the observed behavior and supporting evidence.
- Identify the model, application, or infrastructure decision that enabled it.
- Apply controls at the appropriate layers.
- Re-evaluate the original behavior under comparable conditions.
- Investigate related risks and new behavior introduced by the change.
- Record the result against the tested system version and coverage scope.
Bluejay connects reconnaissance, realistic Digital Humans, adaptive experimentation, structured evaluation, and remediation within this process. Every interaction is expected to contribute something useful: evidence of a failure, evidence that a control held, clarification of an uncertainty, or information that improves the next experiment.
The result is not simply a collection of adversarial conversations. It is an evolving, evidence-based model of how a conversational system behaves under pressure—and a practical method for making that system more secure.
