Modern Mustard Seed
The scoreboard

How the office is actually doing.

We test every specialist against fixed cases with known right answers, and every standing routine checks in when it runs. Sarah marks each report keep or toss. This page reads those records live, every five minutes. When a number is missing, the record does not exist yet, so we show nothing rather than guess.

eval pass rate
Not yet
No sweep reported
specialists tested
Not yet
No sweep reported
routine runs, 14 days
Not yet
No check-ins reported
reports kept
Not yet
No verdicts yet

The scoreboard feed is not reporting yet (feed answered 404). This page fills in on its own once it does.

The evals

Graded against known answers.

Each specialist gets the same cases every sweep: a proposal that must land on the right package, a page with a planted wrong price, an email that must not promise what we do not sell. A change to a specialist’s rules reruns its cases before it reaches real work. Weakest first, so the work to do is at the top.

No eval sweep has reported yet.

The first sweep posts here the moment it finishes. Until then there is nothing honest to show.

The standing routines

Who showed up for work.

Some of our specialists work on a clock, most of them before anyone is awake. Each one checks in when it starts and again when its report is written. A square is one day, oldest on the left.

No routine has checked in yet.

Each routine reports here the first time it runs on its clock.

Why we publish it

Agents you can check on.

Anyone can say their AI works. We would rather show the record, including the misses, because a miss on this page is a fix on our list.

The cases live next to the specialists they test. A change to a specialist’s rulebook reruns its cases before it touches real work, so a rule that makes one worse never ships quietly.

Get an office like this