How the office is actually doing.
We test every specialist against fixed cases with known right answers, and every standing routine checks in when it runs. Sarah marks each report keep or toss. This page reads those records live, every five minutes. When a number is missing, the record does not exist yet, so we show nothing rather than guess.
- eval pass rate
- Not yet No sweep reported
- specialists tested
- Not yet No sweep reported
- routine runs, 14 days
- Not yet No check-ins reported
- reports kept
- Not yet No verdicts yet
The scoreboard feed is not reporting yet (feed answered 404). This page fills in on its own once it does.
Graded against known answers.
Each specialist gets the same cases every sweep: a proposal that must land on the right package, a page with a planted wrong price, an email that must not promise what we do not sell. A change to a specialist’s rules reruns its cases before it reaches real work. Weakest first, so the work to do is at the top.
The first sweep posts here the moment it finishes. Until then there is nothing honest to show.
Who showed up for work.
Some of our specialists work on a clock, most of them before anyone is awake. Each one checks in when it starts and again when its report is written. A square is one day, oldest on the left.
Each routine reports here the first time it runs on its clock.
Agents you can check on.
Anyone can say their AI works. We would rather show the record, including the misses, because a miss on this page is a fix on our list.
The cases live next to the specialists they test. A change to a specialist’s rulebook reruns its cases before it touches real work, so a rule that makes one worse never ships quietly.
Get an office like this