Receipts, not claims. Method, current numbers, and limits.
The documented failure of AI feedback is flattery. Public tests have caught AI feedback tools rating amateur scripts above produced classics, and the Writers Guild's own member guidance names sycophancy as one of the known hazards of AI tools. A reader that praises everything is worse than no reader at all: it costs you rewrites you needed and confidence you did not earn.
So before Auteurium's AI readers were allowed to read anyone's pages, we built them an exam. It is the same idea as a flight simulator: scenes with known, deliberately planted problems, and clean scenes with none, run again and again to see what the readers actually do.
WGA member guidance naming sycophancy: wga.org, AI: know your rights.
Bench of 15 curated scenes (12 planted-flaw, 3 clean), scored across repeated trials. As of July 2026.
This is a small bench by design. It is curated, not crowd-sourced: every scene, every planted flaw, and every scoring judgment was made by hand, and one fixture is currently pulled from scoring while we repair its design. Fifteen scenes cannot prove our readers are always right. What the bench does prove is narrower and more useful: these readers find real planted problems, they do not invent problems on clean pages, and every change we make to them is measured against that standard.
Every calibration change to the readers is scored against the bench before it ships. If catch, churn, or articulation gets worse, the change is reverted, whatever we hoped it would do. This has already killed changes we liked. The bench outranks our taste, and it definitely outranks the readers'.
If the room reads your scene and finds nothing worth your time, it says so and stops. That is not a malfunction. It is the entire point: a reader you can trust when it is quiet is the only kind worth listening to when it is not.