Back to blog

The last signature

23 September 2026· 5 min readaiethicspolicyclinicians
The last signature

On 15 September, a San Francisco startup called TypeSafe AI came out of two years of stealth with its first model, Jev. By the end of that week its launch post had gathered nearly 1,900 points on Hacker News, integration pull requests were stacking up across open-source projects, and Vercel's CEO had already wired it in as the safety reviewer for every agent command: up to 18x faster, he reported, and more accurate.

Jev is easy to explain because it does so little, on purpose. It doesn't write essays. It answers multiple-choice questions. State goes in, and a typed decision comes back with a probability attached, in about a tenth of a second, for a price that rounds to pocket change. You set the threshold at which it acts alone and the threshold at which it asks a human. One of its advertised jobs is checking other AI: scoring, verifying and catching jailbreaks. The verifier's role, finally cheap enough to leave running everywhere.

Healthcare has wanted exactly this: cheap, instant checking. On 17 September, Robert Scoble posted that his whole team at Sully.ai, a hospital-automation startup, was building with Jev less than 48 hours after launch. I replied the next day guessing a healthcare example would show up fast. Yohei Nakajima, a VC and builder, answered that hospital was the better showcase for what this kind of AI can actually do. He wasn't wrong: triage routing, prior authorisation, monitoring, a first read on a mole before it reaches a dermatologist, are exactly the bounded decisions where speed pays. Imagine every one of those decisions checked instantly, and nobody able to say who owns the ones that go wrong.

The engine and its brakes

The general argument arrived in February, in an economics paper from MIT, WashU and UCLA: "the binding constraint on growth is no longer intelligence but human verification bandwidth". Christian Catalini, its lead author, adds a blunter image in his follow-up essay: we've built an engine that scales faster than its brakes. Automating work gets exponentially cheaper, while verifying it stays tied to human time and expertise, and to the speed at which the real world reveals mistakes.

Meanwhile healthcare demand is compounding on two fronts: sick-care for an ageing population and pro-active care for a growing group chasing a longer healthspan. Automation absorbs that, and a model that never tires of cheap multiple-choice questions is the obvious absorber.

A model that always returns well-formed decisions can still be confidently wrong inside the form: that's the trap. "Zero hallucinations" means zero malformed outputs, not zero errors. A malformed answer fails loudly; a well-formed wrong answer, routed automatically in a tenth of a second, fails silently and at volume. In software that's a support ticket. In a clinic it's a patient.

The NHS is already living the quiet version. At the end of August, Healthwatch England reported that an AI scribe had recorded a woman as having demyelination, serious nerve damage, when her MRI result was "null demyelination". She was an NHS health professional herself, and she caught it. The same report describes patients identifying drug mix-ups and a dropped repeat-prescription instruction that clinicians hadn't noticed. GPs and hospital doctors in England already use 27 different scribes; the regulator hasn't classified them as medical devices. The verification layer meant to be the clinician keeps turning out to be the patient.

Who owns the threshold

The community use-case guide for Jev is blunt that it's "not a replacement for authorization, deterministic validation, policy enforcement, or human review in high-impact workflows". Good. But somebody still sets the threshold where the machine acts alone, and that threshold is a governance object, not a config setting. Calibration is self-reported, patient populations drift, and a 95% line tuned on last year's data passes one case in twenty without anyone seeing it.

The founders named the model after William Stanley Jevons. Jevons noticed that more efficient steam engines didn't reduce coal consumption; efficiency made coal cheap enough to burn everywhere, so Britain burned more of it. Cheap checking travels the same road: every drop in the cost of a decision invites more decisions. The volume of machine-judged calls grows; the number of humans able to underwrite them doesn't. The tail grows with the volume.

And the chain, wherever it ends, ends at a name. US malpractice law already says so: clinicians hold a nondelegable duty of care, so responsibility for AI-assisted decisions rests with the physician, however many layers of software did the checking first.

AI checks AI

  • Throughput: millions of checks an hour
  • Cost: fractions of a cent
  • Failure mode: well-formed, silently wrong
  • What it scales: execution

A named human still signs

  • Throughput: a few decisions a day
  • Cost: attention, licence, reputation
  • Failure mode: visible, interruptible, explicable
  • What it scales: accountability

In August I argued that every AI agent needs a register, a named owner and a way to be withdrawn. Jev sharpens that economics rather than dissolving it: the register is where the scarce resource lives.

So the budget line I'd defend in the next board meeting is not more agents. It's verification capacity, audit trails and a small number of people whose job is to be the accountable name.

Which AI decisions already flow through your department without a name under them? List them this week. For each, write down who signs. Any blank row is next year's incident report.

💥 May this inspire you to count your named humans before you count your agents.