Methodology & Findings
A qualitative, workshop-driven study using hybrid affinity mapping to surface patterns across AI red teaming competitors.
RQ1 — Data Visualization
What type of data visualization are used and are they effective?
RQ2 — Actionability
How do competitors support users in acting on the findings?
Data Sources
2 Workshop Transcripts
2×2 Plot
How Might We (HMW) Ideation
Analytical Pipeline
Step 1
Transcript Synthesis
Workshop transcripts were reviewed and synthesized.
Transcripts
▸
Step 2
Deductive Coding
Notes were mapped to 5 pre-defined high-level categories.
Actionability & Reporting
Comprehension / Cognitive Load
Navigation Flow
Data Visualization
Features / Elements
▸
Step 3
Inductive Coding
Iterative refinement of data mapping and themes.
▸
Step 4
Standout Flagging
Identified positively perceived standout features.
▸
Step 5
Convergence & Divergence
Other data sources used to identify convergence and divergence with transcript findings.
2×2 Plot
HMW Board
Limitations
This review covers a small pool of 6 competitors and is not intended to represent the full landscape of AI red teaming products. Analysis was also constrained to publicly available UIs, which may not reflect the complete functionality of each tool. These factors limit the generalizability of the findings.
RQ1 — Data Visualization Findings
Competitors prioritize data richness over delivery.
Across all 6 competitors, the common approach is to show everything — resulting in a UI that overwhelms rather than informing users. This makes it difficult for users to understand the vulnerability status of their models and what to look at first.
Critical
📋
Visual and cognitive overload
Dense layouts force users to work to find insight — rather than surfacing it.
Cited in 2 of 6 competitors
High
📊
Overly complex diagram types
Diagrams that demand interpretation instead of delivering insight fail their users.
5 of 6 found Sankey diagrams unreadable
Moderate
❓
Unclear scoring logic
If users question the scoring mechanism, they can't act on it.
Cited in 2 of 6 competitors
"I probably spend as much time trying to understand what they're trying to tell me as I would doing it myself."
— Workshop Participant
Core insight
Effective data visualization isn't about showing all the findings — it's about helping users understand the status of their models at one glance and the risks clearly.
RQ2 — Actionability
Users are left to connect the dots themselves.
Even when findings are surfaced, the path from "what was found" to "what do I do next" is unclear. Users are unsure what vulnerability needs to be prioritized, and how a vulnerability impacts the business.
Critical
🏢
Missing business context
The "why should I care" layer is almost entirely absent across competitors.
Gap identified across 5 of 6 competitors
Critical
⚖️
No clear prioritization signal
When everything surfaces at the same visual weight, nothing feels urgent.
Cited by a participant
Moderate
📝
Recommendations restate findings
Users need specific next steps, not a restatement of the problem.
Observed in 2 of 6 competitors
"How do I connect the dots between what I wanted to be tested and the results they're showing me?"
— Workshop Participant
"Understanding why I should care about the results is the key missing component from almost everyone. I'd have no idea how to interpret results."
— Workshop Participant
Core insight
Typically, the assessment becomes a dead-end rather than a starting point for fixing issues.