All case studies

Project 01

Generative AI in Healthcare

A structured risk assessment for a hypothetical generative-AI assistant in a healthcare organisation.

Decision context

This independent case study considers a hypothetical generative-AI assistant that helps non-clinical staff retrieve approved healthcare policies and draft internal summaries. It is not intended to diagnose, prescribe, triage or replace professional judgement.

The initial governance question is not “Does the model work?” It is: Can the organisation define a narrow, beneficial use, understand the ways it could fail and keep accountable humans in control?

Intended use and boundaries

  • Retrieve content only from an approved, version-controlled policy library.
  • Produce drafts that require human review before use or distribution.
  • Refuse clinical advice, patient-specific recommendations and emergency guidance.
  • Prohibit entry of identifiable patient data unless a separately approved design supports it.
  • Log sources, user feedback, incidents and overrides without collecting unnecessary content.

Stakeholders and potential harms

| Stakeholder | Foreseeable harm | Risk driver | |---|---|---| | Patients and service users | Incorrect information influences a care-related decision | Hallucination, outdated source, automation bias | | Staff users | Over-reliance or loss of confidence | Poor calibration, weak training, unclear limitations | | Information governance | Personal or sensitive data exposure | Prompt content, retention, supplier access | | Clinical safety leadership | Safety signal missed or routed too slowly | Inadequate monitoring and incident thresholds | | Organisation | Regulatory, financial and reputational harm | Weak ownership, evidence or vendor controls |

Inherent risk view

The scenario starts with high inherent risk because healthcare information can influence safety-critical activity, even where the tool is nominally administrative. Select cells below to explore how likelihood and impact combine; the example default of likelihood 3 and impact 4 produces a high score of 12.

Interactive control

5 × 5 risk matrix

Selected 12High
Likelihood →
Impact →
L LowM ModerateH HighC Critical

Control design

Preventive controls

  • Restrict retrieval to approved, current sources with named content owners.
  • Apply role-based access, data minimisation and prompt-level sensitive-data warnings.
  • Test refusal behaviour for clinical advice, emergency use and out-of-scope requests.
  • Require documented human approval for any output that informs operational or care decisions.
  • Complete supplier security, privacy, resilience and subcontractor due diligence.

Detective controls

  • Display source links and measure unsupported-claim rates in a representative test set.
  • Monitor override patterns, repeated refusals, sensitive-data attempts and user feedback.
  • Sample outputs for accuracy, harmful bias, outdated content and evidence quality.
  • Define thresholds that trigger restriction, rollback, investigation or suspension.

Corrective controls

  • Maintain a clear AI incident route linked to clinical safety, information governance and cyber response.
  • Support rapid source withdrawal, model rollback and user notification.
  • Record root cause, affected decisions, control failure and lessons learned.

Human oversight design

Human oversight is meaningful only if the reviewer has competence, time, authority, information and an escalation path. A generic “human in the loop” statement is therefore insufficient. Reviewers should see the source, understand limitations, be able to reject the output and know when the system must not be used.

Residual risk and decision

The proposed controls could reduce likelihood, but the potential impact of an unsafe healthcare-related output remains material. The recommended decision is a limited, monitored pilot for non-clinical policy retrieval only, subject to privacy review, safety sign-off, representative testing and explicit stop criteria. Expansion into patient-specific or clinical decision support would require a new assessment and specialist clinical-safety governance.

Monitoring evidence

  • Unsupported-claim and source-citation rates.
  • Out-of-scope request and refusal performance.
  • Sensitive-data events and near misses.
  • User overrides, complaints and incident trends.
  • Content freshness and withdrawn-source response time.
  • Control-owner attestations and action closure.

This demonstrates a risk decision—not a claim that the hypothetical system is safe by default.