AUTEURIUM.

How our readers are tested

Receipts, not claims. Method, current numbers, and limits.

The problem this solves

The documented failure of AI feedback is flattery. Public tests have caught AI feedback tools rating amateur scripts above produced classics, and the Writers Guild's own member guidance names sycophancy as one of the known hazards of AI tools. A reader that praises everything is worse than no reader at all: it costs you rewrites you needed and confidence you did not earn.

So before Auteurium's AI readers were allowed to read anyone's pages, we built them an exam. It is the same idea as a flight simulator: scenes with known, deliberately planted problems, and clean scenes with none, run again and again to see what the readers actually do.

WGA member guidance naming sycophancy: wga.org, AI: know your rights.

The method

The numbers, honestly framed

11 / 11
planted flaws caught in the current scored bench
1
net churn note across the clean-scene probes
14 / 14
articulation marks: notes name the real problem

Bench of 15 curated scenes (12 planted-flaw, 3 clean), scored across repeated trials. As of July 2026.

This is a small bench by design. It is curated, not crowd-sourced: every scene, every planted flaw, and every scoring judgment was made by hand, and one fixture is currently pulled from scoring while we repair its design. Fifteen scenes cannot prove our readers are always right. What the bench does prove is narrower and more useful: these readers find real planted problems, they do not invent problems on clean pages, and every change we make to them is measured against that standard.

The keep-or-revert law

Every calibration change to the readers is scored against the bench before it ships. If catch, churn, or articulation gets worse, the change is reverted, whatever we hoped it would do. This has already killed changes we liked. The bench outranks our taste, and it definitely outranks the readers'.

What we deliberately do not publish: the test scenes themselves, the planted flaws, or anything that would let a model be tuned to the exam instead of to the craft.

The limits

Zero notes is a valid read

If the room reads your scene and finds nothing worth your time, it says so and stops. That is not a malfunction. It is the entire point: a reader you can trust when it is quiet is the only kind worth listening to when it is not.

Auteurium ยท Where we stand on AI