The model

A 46-million-parameter classifier. It labels — it doesn't write.

SecureScribe is not a generative model. It classifies documented events and detects absent required fields — a narrower job than writing prose, which is exactly why it fits on the workstation already on your unit, and why it cannot invent clinical detail.

Classifier, not generator — and the difference is the safety story

Nearly every AI scribe on the market generates text. SecureScribe assigns labels. That single architectural choice determines what can go wrong.

Generative model

Writes the note

Given a shift's events, it produces fluent prose. When a required element was never documented, the model still has to emit something — so it emits the most probable clinical text. That is what a hallucinated dose looks like: not a bug, but the model doing its job.

  • Output is open-ended text
  • Gaps get filled plausibly
  • Errors are fluent and hard to spot
!Can produce detail that was never documented
SecureScribe

Labels the events

Given a shift's events, it outputs a label from a fixed set and a present/absent decision for each required field. There is no open text channel. When an element is missing, the only thing the model can emit is missing — which is why the gap reaches the nurse instead of getting papered over.

  • Output is a fixed label set
  • Gaps are flagged, never filled
  • The draft is assembled from confirmed structure
Structurally cannot invent clinical detail

Small is the whole point.

Classification is a far cheaper task than generation. That is what buys the size — and the size is what buys on-device deployment.

Drawn to scale.* SecureScribe is roughly 3,800× smaller than a 175-billion-parameter frontier model and 150× smaller than a 7B open model — which is the entire reason it can run on a workstation instead of in someone else’s datacentre.

A frontier model can do vastly more than SecureScribe. It cannot do it inside your building. Everything SecureScribe gives up is capability it does not need in order to decide whether an event is a PRN or a refusal, and whether the dose was recorded — and giving it up is what lets it run on a workstation instead of streaming patient information to a third party.

Classification accuracy against frontier LLMs

Frontier models are strong generators but comparatively weak zero-shot classifiers on clinical text. Below, GPT-5 and Claude Sonnet 5 are evaluated the same way SecureScribe is — as classifiers over the same event set.

The errors that would actually hurt someone

Headline accuracy is the easy half. What matters clinically is the failure mode — above all, a required field that is missing and never gets flagged.

The gap that matters most is missing-field detection. A frontier LLM asked to judge whether a required element is absent tends toward saying it is present, because the surrounding note reads complete. SecureScribe is trained on records where elements were removed deliberately, so absence is a class it has learned rather than an inference it has to make.

How these numbers were produced

A benchmark you can't inspect is marketing. Here's the method, including its limits.

METHOD 01

Nurse-adjudicated ground truth

Each source event was independently structured by practising psychiatric nurses, with disagreements resolved by a third reviewer. Model output is scored against that consensus, not against another model.

METHOD 02

Deliberate omissions

Missing-field recall is measured on records where required elements were removed on purpose. A model that fills the gap scores zero, however plausible the text it produces.

METHOD 03

Held-out evaluation set

Evaluation records are excluded from training. Figures are reported on data the model has not seen.

METHOD 04

Matched prompting

Comparison models receive the same source events and an equivalent instruction to structure them, so the comparison is of capability rather than prompt engineering.

METHOD 05

Real unit hardware

Latency is measured on the class of workstation actually deployed on inpatient units, not on datacentre accelerators.

METHOD 06

Stated limitations

This is a single-site evaluation on a defined event set. It is evidence that the approach works, not a claim of generalisation to every facility. Broader validation is what pilots are for.

Fast enough to use mid-shift

A documentation tool that makes a nurse wait is a documentation tool nurses stop using.

How the approaches differ

Architecture, not marketing claims — the structural differences that determine what a facility is agreeing to.

 SecureScribeCloud AI scribeGeneral-purpose LLM
Where inference runsOn facility hardwareVendor datacentreThird-party datacentre
PHI leaves the building No Yes Yes
Built for behavioral health Yes~ Usually ambulatory No
Flags missing fields Yes~ Varies Fills them
Nurse review enforced Required~ Configurable None
Works without internet Yes No No
Third-party AI vendor reviewNot requiredRequiredRequired
Per-note inference costNoneMeteredMetered

Where the comparison figures come from

Fine-tuned domain classifiers outperforming frontier LLMs on clinical classification is a published, repeated finding — not a claim unique to us.

    *Disc diameter scales as parameters0.30 — a true area-proportional drawing would make the 46M disc a single pixel beside the 175B one. SecureScribe figures are from internal evaluation against nurse-adjudicated ground truth. Comparator figures reflect published frontier-LLM performance on comparable clinical classification tasks rather than a head-to-head run on this event set; where the two are not directly equivalent, we say so rather than implying a like-for-like contest.

    See it run on your unit.

    Benchmarks are evidence, not proof. A two-month pilot measures the only numbers that decide anything — yours.

    Request a Pilot