Physician-led adversarial audits that uncover clinical reasoning failures and safety blind spots before your buyers find them in public. Protect your sales pipeline and scale with confidence.
"Dr. Rameesha brings a rare level of rigor to clinical AI evaluation. Her adversarial testing surfaced real execution-boundary failures—cases where insufficient clinical state still resolved to actionable outputs—and did so in a structured, reproducible way. What stands out is her ability to go beyond model performance and identify where systems produce outputs that shouldn’t exist given what they can actually know. That level of clarity is critical in clinical environments. Her work directly strengthened execution gating behavior, helping move from detection to true fail-closed enforcement."Tim Zlomke — Founder, SolaceMedAI
Dr. Chohan serves on the CAVAL™ Global Clinical Advisory Panel at Kairoc Systems Inc., contributing clinical governance and adversarial testing expertise to advance safe, auditable AI infrastructure for healthcare.
We stress-test your system against the complexity of real clinical practice, not clean benchmark data. Every audit runs through READS, our proprietary five-pillar framework:
Diagnostic accuracy, guideline alignment and logical consistency of clinical justification chains.
Demographic invariance and linguistic fairness across dialects, socio-economic signals and communication styles.
Resilience against chaotic shorthand, syntactic negations, narrative distraction and active manipulation.
State tracking, conversational memory, follow-up efficiency and contradiction handling.
Compliance against clinical contraindications, scope-of-practice boundaries and unauthorized PHI exposure.
Instant fail-safe triggers. A system fails the audit immediately if it misses a life-threatening rule-out, shifts output on a demographic swap or leaks PHI — regardless of overall score.
Structured multi-run testing matrices, up to 750 targeted permutations per audit, to eliminate statistical noise and verify reproducibility.
Granular mapping of exactly where justification chains fracture, how prompt drift occurs and why safety boundaries degrade under pressure.
Failure modes mapped to the 7 normative characteristics of the NIST AI Risk Managmenet Framework and translated into financial loss vectors via the Huwyler Threat Taxonomy.
Every audit follows the same disciplined arc — from mapping your model's risk surface to a governance-ready report. Scope, timeline, and investment are set together on a scoping call, calibrated to your system's autonomy tier.
A technical and clinical deep-dive to map out your model's architecture and specific risk vectors.
Bespoke stress-testing of reasoning layers against edge cases, clinical biases, and safety guardrails.
A comprehensive vulnerability ledger and clinical mitigation strategy ready for regulatory review.
We provide risk visibility, not fractional CMO/CSO coverage. The client retains sole accountability for deployment and clinical outcomes.
We evaluate model reasoning, safety guardrails, and adversarial vulnerabilities — not cloud infrastructure, penetration testing, or IT certification.
This is an empirical behavioral stress-test, not a compliance stamp or legal safety guarantee.
Engineering tracks uptime and syntax. Data science tracks accuracy on clean data. Neither catches a model breaking under real clinical friction.
Your clinical board knows what the AI should do. We're trained to find what it can be forced to do.
Internal teams test for intended use. We test for misuse — that's the gap that gets missed internally.
One missed rule-out matters more than a high overall score. We isolate catastrophic failures first, independent of aggregate accuracy.
Hospital networks and institutional buyers want third-party validation. An independent audit shifts your profile from internal negligence to proactive due diligence.
Catching faults early protects your sales pipeline, shortens time-to-market, and prevents public failures that cost investor and buyer trust.
Founder, GarrisonLabs
Most AI safety consultants have never held a stethoscope. I have.
I left clinical practice to found GarrisonLabs and build the READS Framework after watching advanced AI tools fracture under the unpredictable reality of real-world patient care. Today, I pair frontline clinical experience with practical tech skills to help HealthTech founders uncover hidden reasoning failures before they reach a patient or a regulator. Building in this space is high-stakes — you shouldn't have to navigate clinical liability alone.
We don't grade models on a curve. Look inside our independent adversarial audits of industry-leading clinical AI platforms.
A physician-led adversarial audit evaluating Doctronic's clinical reasoning integrity, safety guardrails, and dialogue coherence across a structured matrix of complex renal presentations and life-threatening adjacent diagnoses. The system operates as a patient-facing clinical AI — the highest-stakes autonomy tier for diagnostic output.
READS AUDIT MATRIX: 12 Adversarial Cases Across a Single Disease Domain
Doctronic.ai was subjected to a rigorous 12-case adversarial clinical audit using the READS framework to evaluate its operational boundaries within the Renal Stones domain — specifically chosen because its classic presentation creates strong diagnostic anchoring, while adjacent conditions carry extreme, time-sensitive lethality. The audit exposed a critical performance boundary: while Doctronic scored perfectly on Dialogue & Workflow (100%), its core Clinical Reasoning layer experienced a total collapse (0%), failing to protect against critical rule-outs or resist dangerous patient anchoring.
Standard engineering evaluations routinely validate conversational mechanics, data retention, and linguistic variations. Under these parameters, Doctronic performs exceptionally well. However, when faced with complex clinical overlays — such as atypical geriatric presentations or anatomic modifiers — Doctronic's structural reasoning layer fractures, generating severe logic contradictions and unsafe triage recommendations.
CRITICAL failure in Valid & Reliable metrics. Accountable, Transparent, and Explainable characteristics passed — but systematic diagnostic omissions triggered a system-level failure flag.
All primary reasoning failures map to Unreliable Outputs and Biases domains, carrying high-severity risk scores.
Immediate Integrity, Legal, and Reputation liabilities. Institutional buyers will kill B2B contracts if an independent audit uncovers branch-level triage failures exposing the vendor to malpractice liability.
A variant stress-test audit examining Symptomate's emergency triage reliability versus its underlying diagnostic precision layer. The system's safety floor performed correctly — but the reasoning model behind it failed to evaluate or rule out a life-threatening differential in every single variant tested.
VARIANT STRESS-TEST: Acute Appendicitis Matrix (Female, 35 Years) — Infermedica v6.17.0
Symptomate was evaluated using a single-case variant stress-test focused on Acute Appendicitis in a 35-year-old female. The audit demonstrated that while the platform's safety floor functioned flawlessly — correctly routing 100% of variants to emergency care — the underlying diagnostic precision layer collapsed, exhibiting rigid linguistic anchoring and a complete failure to evaluate secondary rule-outs.
When a triage model correctly triggers an emergency alert, internal engineering teams often assume their safety guardrails are bulletproof. However, if the underlying reasoning model anchors onto an incorrect or anatomically impossible diagnosis while giving that emergency care recommendation, it creates an unearned "confidence label" — generating massive operational workflow confusion and a high liability profile when integrating with live healthcare records.
Latent Logic Faults & Biases (Proxy Discrimination). The system gives founders a false sense of security by masking severe diagnostic errors behind an accurate triage label.
System highly vulnerable to conversational edge cases that bypass triage guardrails. Unearned confidence labels create institutional liability during EMR integration.
A system recommending appendicitis evaluation for a patient who has no appendix demonstrates a foundational flaw in context integration — a procurement-killing discovery during pilot evaluation.
A structured adversarial protocol probing the boundary between Ada's surface-level pattern recognition and deep clinical reasoning. The system performs reliably on textbook single-system presentations — but fails systematically when cases require comorbidity integration, geographic context, or logical rule-out sequencing.
READS AUDIT MATRIX: 11 Adversarial Cases Across 5 Failure Modes
Ada Health's symptom checker was subjected to a structured adversarial protocol probing the operational boundary between surface-level pattern recognition and deep clinical reasoning. The audit identified a categorical performance boundary: Ada performs reliably on textbook, single-system presentations requiring static pattern matching alone, but fails systematically when cases require comorbidity integration, geographic context, or logical questioning sequences.
Under sterile conditions, clinical models achieve near-perfect scores. However, real-world patients present with mixed clinical signals, historical comorbidities, and geographic variables. When evaluated against these layered complexities, the model's structural clinical reasoning layer collapses while its superficial pattern-matching engine continues to run — producing confident-sounding outputs that are clinically dangerous.
Unreliable Outputs — Logic & Factual Hallucination. Sub-optimal clinical reasoning on high-stakes cases exposes healthtech vendors to severe liability and user mistrust.
Institutional buyers and hospital risk review committees do not buy brittle clinical agents. Branch-level collapses discovered during pilot deployment kill commercial contracts instantly.
If an enterprise client's validation team uncovers these failures during a live demo, the commercial contract is dead. Finding them first — in private — is the only viable strategy.
These are the failures we find in independent testing. Imagine what we'll find in yours.
Schedule a Scoping Call →Ready to identify your system's vulnerabilities in a secure, sandboxed environment? Reach out below to schedule a targeted tier or request a custom scoping proposal.
All inquiries are treated with strict confidentiality. Mutual NDAs signed prior to any system access.
Prefer to reach out directly? Email or LinkedIn
To request an ongoing retainer, please contact directly at Email. Retainers are not bookable through the intake portal.