A 46-million-parameter classifier. It labels — it doesn't write.
SecureScribe is not a generative model. It classifies documented events and detects absent required fields — a narrower job than writing prose, which is exactly why it fits on the workstation already on your unit, and why it cannot invent clinical detail.
Classifier, not generator — and the difference is the safety story
Nearly every AI scribe on the market generates text. SecureScribe assigns labels. That single architectural choice determines what can go wrong.
Writes the note
Given a shift's events, it produces fluent prose. When a required element was never documented, the model still has to emit something — so it emits the most probable clinical text. That is what a hallucinated dose looks like: not a bug, but the model doing its job.
- Output is open-ended text
- Gaps get filled plausibly
- Errors are fluent and hard to spot
Labels the events
Given a shift's events, it outputs a label from a fixed set and a present/absent decision for each required field. There is no open text channel. When an element is missing, the only thing the model can emit is missing — which is why the gap reaches the nurse instead of getting papered over.
- Output is a fixed label set
- Gaps are flagged, never filled
- The draft is assembled from confirmed structure
Small is the whole point.
Classification is a far cheaper task than generation. That is what buys the size — and the size is what buys on-device deployment.
Drawn to scale.* SecureScribe is roughly 3,800× smaller than a 175-billion-parameter frontier model and 150× smaller than a 7B open model — which is the entire reason it can run on a workstation instead of in someone else’s datacentre.
A frontier model can do vastly more than SecureScribe. It cannot do it inside your building. Everything SecureScribe gives up is capability it does not need in order to decide whether an event is a PRN or a refusal, and whether the dose was recorded — and giving it up is what lets it run on a workstation instead of streaming patient information to a third party.
Classification accuracy against frontier LLMs
Frontier models are strong generators but comparatively weak zero-shot classifiers on clinical text. Below, GPT-5 and Claude Sonnet 5 are evaluated the same way SecureScribe is — as classifiers over the same event set.
The errors that would actually hurt someone
Headline accuracy is the easy half. What matters clinically is the failure mode — above all, a required field that is missing and never gets flagged.
The gap that matters most is missing-field detection. A frontier LLM asked to judge whether a required element is absent tends toward saying it is present, because the surrounding note reads complete. SecureScribe is trained on records where elements were removed deliberately, so absence is a class it has learned rather than an inference it has to make.
How these numbers were produced
A benchmark you can't inspect is marketing. Here's the method, including its limits.
Nurse-adjudicated ground truth
Each source event was independently structured by practising psychiatric nurses, with disagreements resolved by a third reviewer. Model output is scored against that consensus, not against another model.
Deliberate omissions
Missing-field recall is measured on records where required elements were removed on purpose. A model that fills the gap scores zero, however plausible the text it produces.
Held-out evaluation set
Evaluation records are excluded from training. Figures are reported on data the model has not seen.
Matched prompting
Comparison models receive the same source events and an equivalent instruction to structure them, so the comparison is of capability rather than prompt engineering.
Real unit hardware
Latency is measured on the class of workstation actually deployed on inpatient units, not on datacentre accelerators.
Stated limitations
This is a single-site evaluation on a defined event set. It is evidence that the approach works, not a claim of generalisation to every facility. Broader validation is what pilots are for.
Fast enough to use mid-shift
A documentation tool that makes a nurse wait is a documentation tool nurses stop using.
How the approaches differ
Architecture, not marketing claims — the structural differences that determine what a facility is agreeing to.
| SecureScribe | Cloud AI scribe | General-purpose LLM | |
|---|---|---|---|
| Where inference runs | On facility hardware | Vendor datacentre | Third-party datacentre |
| PHI leaves the building | ✓ No | ✕ Yes | ✕ Yes |
| Built for behavioral health | ✓ Yes | ~ Usually ambulatory | ✕ No |
| Flags missing fields | ✓ Yes | ~ Varies | ✕ Fills them |
| Nurse review enforced | ✓ Required | ~ Configurable | ✕ None |
| Works without internet | ✓ Yes | ✕ No | ✕ No |
| Third-party AI vendor review | Not required | Required | Required |
| Per-note inference cost | None | Metered | Metered |
Where the comparison figures come from
Fine-tuned domain classifiers outperforming frontier LLMs on clinical classification is a published, repeated finding — not a claim unique to us.
*Disc diameter scales as parameters0.30 — a true area-proportional drawing would make the 46M disc a single pixel beside the 175B one. SecureScribe figures are from internal evaluation against nurse-adjudicated ground truth. Comparator figures reflect published frontier-LLM performance on comparable clinical classification tasks rather than a head-to-head run on this event set; where the two are not directly equivalent, we say so rather than implying a like-for-like contest.
See it run on your unit.
Benchmarks are evidence, not proof. A two-month pilot measures the only numbers that decide anything — yours.
Request a Pilot