Clinical AI Red-Teaming & Adversarial Evaluation

Your Enterprise Buyers Will Stress-Test Your Clinical AI.
We Do It First.

Physician-led adversarial audits that uncover clinical reasoning failures and safety blind spots before your buyers find them in public. Protect your sales pipeline and scale with confidence.

READS Framework Clinical Reasoning Evaluation Adversarial Stress-Testing NIST AI RMF Mapping Enterprise Risk Translation Safety Guardrail Auditing Huwyler Threat Taxonomy Physician-Led Methodology READS Framework Clinical Reasoning Evaluation Adversarial Stress-Testing NIST AI RMF Mapping Enterprise Risk Translation Safety Guardrail Auditing Huwyler Threat Taxonomy Physician-Led Methodology
Industry Validation
"Dr. Rameesha brings a rare level of rigor to clinical AI evaluation. Her adversarial testing surfaced real execution-boundary failures—cases where insufficient clinical state still resolved to actionable outputs—and did so in a structured, reproducible way. What stands out is her ability to go beyond model performance and identify where systems produce outputs that shouldn’t exist given what they can actually know. That level of clarity is critical in clinical environments. Her work directly strengthened execution gating behavior, helping move from detection to true fail-closed enforcement."
Tim Zlomke — Founder, SolaceMedAI
Clinical Advisory Leadership

Dr. Chohan serves on the CAVAL™ Global Clinical Advisory Panel at Kairoc Systems Inc., contributing clinical governance and adversarial testing expertise to advance safe, auditable AI infrastructure for healthcare.

The Audit Methodology

Our Proprietary READS Framework

We stress-test your system against the complexity of real clinical practice, not clean benchmark data. Every audit runs through READS, our proprietary five-pillar framework:

R

Reasoning Integrity

Diagnostic accuracy, guideline alignment and logical consistency of clinical justification chains.

E

Equity & Fairness

Demographic invariance and linguistic fairness across dialects, socio-economic signals and communication styles.

A

Adversarial Robustness

Resilience against chaotic shorthand, syntactic negations, narrative distraction and active manipulation.

D

Dialogue & Workflow

State tracking, conversational memory, follow-up efficiency and contradiction handling.

S

Safety & Guardrails

Compliance against clinical contraindications, scope-of-practice boundaries and unauthorized PHI exposure.

Disqualifying Rules Conditions (DRCs)

Instant fail-safe triggers. A system fails the audit immediately if it misses a life-threatening rule-out, shifts output on a demographic swap or leaks PHI — regardless of overall score.

Rigor-Calibrated Sampling

Structured multi-run testing matrices, up to 750 targeted permutations per audit, to eliminate statistical noise and verify reproducibility.

Deep-Dive Discrepancy Analysis

Granular mapping of exactly where justification chains fracture, how prompt drift occurs and why safety boundaries degrade under pressure.

Strategic Governance Mapping

Failure modes mapped to the 7 normative characteristics of the NIST AI Risk Managmenet Framework and translated into financial loss vectors via the Huwyler Threat Taxonomy.

Statistical methodology including Wilson Score CIs through GLMM with Crossed Random Effects is calibrated per engagement scope.
How An Engagement Runs

The Engagement Process

Every audit follows the same disciplined arc — from mapping your model's risk surface to a governance-ready report. Scope, timeline, and investment are set together on a scoping call, calibrated to your system's autonomy tier.

01

Discovery & Scoping

A technical and clinical deep-dive to map out your model's architecture and specific risk vectors.

02

Adversarial Probing

Bespoke stress-testing of reasoning layers against edge cases, clinical biases, and safety guardrails.

03

Governance & Reporting

A comprehensive vulnerability ledger and clinical mitigation strategy ready for regulatory review.

Request a Scoping Call →
Inquire About Clinical Evaluation Availability →
Operational Clarity

Scope & Boundaries

Not Provided

No Clinical Sign-Off

We provide risk visibility, not fractional CMO/CSO coverage. The client retains sole accountability for deployment and clinical outcomes.

Not Provided

No Technical Security Audit

We evaluate model reasoning, safety guardrails, and adversarial vulnerabilities — not cloud infrastructure, penetration testing, or IT certification.

Not Provided

Not a Regulatory Certification

This is an empirical behavioral stress-test, not a compliance stamp or legal safety guarantee.

Why External Audit

Why An External Audit

Standard QA Misses Clinical Collapse

Engineering tracks uptime and syntax. Data science tracks accuracy on clean data. Neither catches a model breaking under real clinical friction.

Internal Advisors Guide; We Break

Your clinical board knows what the AI should do. We're trained to find what it can be forced to do.

Builder's Blindness Is Real

Internal teams test for intended use. We test for misuse — that's the gap that gets missed internally.

Zero-Tolerance Risk Detection

One missed rule-out matters more than a high overall score. We isolate catastrophic failures first, independent of aggregate accuracy.

The Enterprise Sales Buffer

Hospital networks and institutional buyers want third-party validation. An independent audit shifts your profile from internal negligence to proactive due diligence.

Commercial & Regulatory De-risking

Catching faults early protects your sales pipeline, shortens time-to-market, and prevents public failures that cost investor and buyer trust.

FAQ

Common Questions

Do you provide clinical safety certifications? +
No. We provide independent, real-time adversarial testing and document how your system behaved under a specific matrix of high-stress scenarios — not compliance stamps or software validations.
Why mixed-effects models (GLMM) for Tier 3? +
GLMM with Crossed Random Effects mathematically isolates clinical complexity from superficial prompt-wording variation, proving to regulators exactly what drives your model's stability.
What do our engineers actually receive? +
Depending on tier: a structured READS evaluation, a systematic Failure Coding Log, actionable engineering recommendations, NIST AI RMF / Huwyler risk mappings, and a deployment roadmap.
How long does an audit take? +
10–14 business days for Exploratory, up to 3–7 weeks for Comprehensive and Regulatory-Grade.
Dr. Rameesha Chohan

Dr. Rameesha Chohan

Founder, GarrisonLabs

◆ Licensed Physician, Clinical AI Auditor & HealthTech Advisor
About

Dr. Rameesha Chohan
Founder, GarrisonLabs

Most AI safety consultants have never held a stethoscope. I have.

I left clinical practice to found GarrisonLabs and build the READS Framework after watching advanced AI tools fracture under the unpredictable reality of real-world patient care. Today, I pair frontline clinical experience with practical tech skills to help HealthTech founders uncover hidden reasoning failures before they reach a patient or a regulator. Building in this space is high-stakes — you shouldn't have to navigate clinical liability alone.

Portfolio — Independent Audits

Empirical Evidence:
Where Clinical AI Breaks.

We don't grade models on a curve. Look inside our independent adversarial audits of industry-leading clinical AI platforms.

Independent Portfolio Audit · May 2026

Doctronic.ai Clinical Assistant

Patient-Facing Clinical AI Full READS Five-Pillar Evaluation Domain: Renal Stones 12 Adversarial Cases

A physician-led adversarial audit evaluating Doctronic's clinical reasoning integrity, safety guardrails, and dialogue coherence across a structured matrix of complex renal presentations and life-threatening adjacent diagnoses. The system operates as a patient-facing clinical AI — the highest-stakes autonomy tier for diagnostic output.

Expand

READS AUDIT MATRIX: 12 Adversarial Cases Across a Single Disease Domain

25.0%
Cases in RED Zone
4.33 / 5
Avg. RED Severity Score
0%
Clinical Reasoning Score (P1)

Executive Summary

Doctronic.ai was subjected to a rigorous 12-case adversarial clinical audit using the READS framework to evaluate its operational boundaries within the Renal Stones domain — specifically chosen because its classic presentation creates strong diagnostic anchoring, while adjacent conditions carry extreme, time-sensitive lethality. The audit exposed a critical performance boundary: while Doctronic scored perfectly on Dialogue & Workflow (100%), its core Clinical Reasoning layer experienced a total collapse (0%), failing to protect against critical rule-outs or resist dangerous patient anchoring.

The Latent Vulnerability

Standard engineering evaluations routinely validate conversational mechanics, data retention, and linguistic variations. Under these parameters, Doctronic performs exceptionally well. However, when faced with complex clinical overlays — such as atypical geriatric presentations or anatomic modifiers — Doctronic's structural reasoning layer fractures, generating severe logic contradictions and unsafe triage recommendations.

Key Findings & Logged Failures

  • Anchoring and Omission Bias: In case RS-P1-001, the system anchored heavily on renal colic and completely failed to list an Abdominal Aortic Aneurysm (AAA) in its differential — completely missing a red-flag triad (older male, mechanical strain, dizziness) that carries an 80%+ mortality rate if missed.
  • Comorbidity and Geriatric Blindness: The system failed to escalate an obstructing stone in a patient with a solitary kidney (RS-P3-002) — a surgical emergency — and misattributed acute infection to musculoskeletal strain in a 72-year-old diabetic woman.
  • Logic Contraindications & Textual Gaslighting: When a patient pushed back on an assessment, the AI actively fabricated historical data ("Your dizziness was a one-time event right after exertion, not ongoing") to defensively protect its initial incorrect reasoning.
  • Dialogue vs. Reasoning Tradeoff: Flawless 100% in Dialogue & Workflow (P5) and 67% in Adversarial Robustness (P3) — but 0% in Clinical Reasoning (P1).

Business & Enterprise Impact

NIST Trustworthiness

CRITICAL failure in Valid & Reliable metrics. Accountable, Transparent, and Explainable characteristics passed — but systematic diagnostic omissions triggered a system-level failure flag.

Huwyler Threat Taxonomy

All primary reasoning failures map to Unreliable Outputs and Biases domains, carrying high-severity risk scores.

CIA-LR Corporate Exposure

Immediate Integrity, Legal, and Reputation liabilities. Institutional buyers will kill B2B contracts if an independent audit uncovers branch-level triage failures exposing the vendor to malpractice liability.

Variant Independent Testing · May 2026

Symptomate Symptom Checker (Infermedica)

Consumer Symptom Checker Variant Stress-Test Protocol Domain: Acute Appendicitis 8 Adversarial Variants

A variant stress-test audit examining Symptomate's emergency triage reliability versus its underlying diagnostic precision layer. The system's safety floor performed correctly — but the reasoning model behind it failed to evaluate or rule out a life-threatening differential in every single variant tested.

Expand

VARIANT STRESS-TEST: Acute Appendicitis Matrix (Female, 35 Years) — Infermedica v6.17.0

8 / 8
Emergency Triage Validated
0 / 8
Ectopic Pregnancy Ruled Out
41%
Avg. Diagnostic Accuracy

Executive Summary

Symptomate was evaluated using a single-case variant stress-test focused on Acute Appendicitis in a 35-year-old female. The audit demonstrated that while the platform's safety floor functioned flawlessly — correctly routing 100% of variants to emergency care — the underlying diagnostic precision layer collapsed, exhibiting rigid linguistic anchoring and a complete failure to evaluate secondary rule-outs.

The Latent Vulnerability

When a triage model correctly triggers an emergency alert, internal engineering teams often assume their safety guardrails are bulletproof. However, if the underlying reasoning model anchors onto an incorrect or anatomically impossible diagnosis while giving that emergency care recommendation, it creates an unearned "confidence label" — generating massive operational workflow confusion and a high liability profile when integrating with live healthcare records.

Key Findings & Logged Failures

  • Complete Rule-Out Failure: 0 out of 8 variants successfully evaluated or explicitly ruled out an ectopic pregnancy, despite the patient profile directly meeting high-risk demographic and clinical criteria.
  • Rigid Linguistic Anchoring: When a variant explicitly stated the patient had a prior appendectomy — making acute appendicitis anatomically impossible — the platform continuously ranked appendicitis as a primary differential because it could not logically override its initial text token anchor. Accuracy dropped to 41%.
  • Safety Floor vs. Diagnostic Precision Gap: 8 of 8 variants correctly triggered "Emergency Attendance" classification — but the diagnostic reasoning behind the recommendation was structurally flawed in every case.

Business & Enterprise Impact

Threat Vector Domain

Latent Logic Faults & Biases (Proxy Discrimination). The system gives founders a false sense of security by masking severe diagnostic errors behind an accurate triage label.

The Exposure

System highly vulnerable to conversational edge cases that bypass triage guardrails. Unearned confidence labels create institutional liability during EMR integration.

The Sales Bottleneck

A system recommending appendicitis evaluation for a patient who has no appendix demonstrates a foundational flaw in context integration — a procurement-killing discovery during pilot evaluation.

Independent Adversarial Audit · April 2026

Ada Health Symptom Checker

Consumer Clinical AI Full READS Five-Pillar Evaluation 5 Failure Mode Categories 11 Adversarial Cases

A structured adversarial protocol probing the boundary between Ada's surface-level pattern recognition and deep clinical reasoning. The system performs reliably on textbook single-system presentations — but fails systematically when cases require comorbidity integration, geographic context, or logical rule-out sequencing.

Expand

READS AUDIT MATRIX: 11 Adversarial Cases Across 5 Failure Modes

45.5%
Cases in RED Zone
9.2 / 10
Avg. RED Severity Score
0 / 5
Danger Diagnoses Ranked

Executive Summary

Ada Health's symptom checker was subjected to a structured adversarial protocol probing the operational boundary between surface-level pattern recognition and deep clinical reasoning. The audit identified a categorical performance boundary: Ada performs reliably on textbook, single-system presentations requiring static pattern matching alone, but fails systematically when cases require comorbidity integration, geographic context, or logical questioning sequences.

The Latent Vulnerability

Under sterile conditions, clinical models achieve near-perfect scores. However, real-world patients present with mixed clinical signals, historical comorbidities, and geographic variables. When evaluated against these layered complexities, the model's structural clinical reasoning layer collapses while its superficial pattern-matching engine continues to run — producing confident-sounding outputs that are clinically dangerous.

Key Findings & Logged Failures

  • Pattern-Matching Boundary: Ada achieved 100% accuracy on classic single-diagnosis baseline presentations — demonstrating strong textbook correlation but confirming the system relies on surface matching, not reasoning.
  • Comorbidity and Context Collapse: 45.5% of high-stakes adversarial cases fell into the critical RED Zone — representing total failure in safe diagnostic generation or severe triage downgrades.
  • Zero Critical Danger Diagnoses Ranked: Across 100% of RED Zone cases, the system completely failed to include or rank the true, time-critical danger diagnosis in its primary differential list.
  • Critical Case Incoherence: Structural inability to integrate geographic endemic risks (e.g., malaria exposure variables) into its core questioning branch — causing severe clinical misdirection on high-acuity cases.

Business & Enterprise Impact

Threat Vector Domain

Unreliable Outputs — Logic & Factual Hallucination. Sub-optimal clinical reasoning on high-stakes cases exposes healthtech vendors to severe liability and user mistrust.

The Exposure

Institutional buyers and hospital risk review committees do not buy brittle clinical agents. Branch-level collapses discovered during pilot deployment kill commercial contracts instantly.

The Sales Bottleneck

If an enterprise client's validation team uncovers these failures during a live demo, the commercial contract is dead. Finding them first — in private — is the only viable strategy.

Ready?

Find Your System's Breaking Point
Before Your Buyers Do.

These are the failures we find in independent testing. Imagine what we'll find in yours.

Schedule a Scoping Call →
Get In Touch

Find Your System's Breaking Point
Before Your Buyers Do.

Ready to identify your system's vulnerabilities in a secure, sandboxed environment? Reach out below to schedule a targeted tier or request a custom scoping proposal.

Request a Scoping Call

All inquiries are treated with strict confidentiality. Mutual NDAs signed prior to any system access.

Prefer to reach out directly? Email or LinkedIn

Ongoing Quarterly Retainer

To request an ongoing retainer, please contact directly at Email. Retainers are not bookable through the intake portal.

What Happens After You Reach Out

1

Pre-Call Strategy Review

Before we meet, I'll personally review your platform's public-facing product and the challenges you outline in this form.

2

Streamlined Scheduling

A calendar invite with secure video conference details arrives within 1–3 business days.

3

Focused Scoping Session

Our 30-minute call establishes exact clinical scope, deployment milestones, and transparent pricing.

Book Directly

Skip the form and schedule a risk assessment consultation directly on my calendar.

Schedule a Call →

Enterprise Confidentiality Guarantee

All initial inquiries, platform descriptions, and scoping communications are treated with absolute confidentiality. Standard mutual NDAs are signed prior to any system environment access or specialized clinical testing.