Field records · measured, not projected

Everything here is a number we produced ourselves.

No client logos we have not earned. No testimonials we did not receive. No industry statistics borrowed from a report. When we have not measured something, we say so and leave the space empty.

This is the shortest page on most agency websites. We would rather it be the longest on ours.

Record 01 · complete

Autonomous email triage, fourteen days on a live inbox.

A scheduled AI agent ran twice a day against a real working inbox for two weeks. It read the body of every message, not just the sender, then filed, flagged, archived, or escalated according to a written policy it also helped build.

It caught things a filter cannot. A phishing reply dressed as a support thread, routing identity verification to an off-brand domain. Two marketing blasts disguised as personal email, one of which was only detectable from a tracking beacon in the body. A single vendor family turned out to use seven sender addresses that needed routing three different ways, which is why sender rules were the wrong abstraction and judgment was the product.

It also had a real gap. 4 of 27 scheduled runs left no record. We checked two: both had genuinely failed — a timeout, a connection error. The other two, we still don’t know. What was missing wasn’t the error itself — it was the alert. On a personal inbox that is a shrug. On a client’s lead notifications it is a lost deal, which is why every build we ship now alerts on absence.

Record 01 · 14-day live trialOur own inbox
Emails dispositioned146
Senders learned automatically54
Phishing attempts flagged1
Disguised bulk mail caught2
False archives of important mail0
Scheduled runs27
Scheduled runs with no record4 of 27
Measured on our own operations before we sold anything to anyone. The gap is on this card on purpose. Two of the four we checked and found real errors nobody was alerted to; the other two we still don’t know. An agency that has never found a failure has never looked for one.
Record 01 · the pipeline itself

This is the system that produced those numbers — including the one that failed.

Below is the actual shape of it. The first state is what we built. The second is what we run now, and the difference between them is the only reason the gap above is on this page at all.

▸ TRIGGERscheduled run
▸ FETCHunread mail
◆ READthe body, not the sender
◆ CLASSIFYcustomer · coupon · phishing
▸ ACTfile · flag · archive · escalate
◇ MONITORalert on absence
4 of 27 runs left no record
and nothing said so

A run can fail before it ever writes a usable error. Every monitor we had was watching the output, so when a run didn’t complete, nothing had an opinion about it. The fix is the node at the bottom, and the direction of its arrow: it watches the trigger, not the result. That is now a build standard on every system we ship.

Field record · liveoperational
Emails dispositioned
146
Senders learned automatically
54
False archives of important mail
0
Scheduled runs with no record — the number nobody else publishes
4/27

Measured on our own operations before we sold anything to anyone. The gap is on this card on purpose. Two of the four we checked were real errors nobody was alerted to; the other two, we still don’t know. An agency that has never found a failure has never looked.

The honest part · what this record does not show

A number without its limits is marketing.

This was our inbox, not a client’s

Around six threads a day. A shared support inbox running two hundred a day is roughly thirty times the volume against the same build. We publish the multiplier rather than a promise, and your audit produces your real number.

Fourteen days is a trial, not a track record

Long enough to find a silent failure mode. Not long enough to claim reliability. Record 02 is the harder test: somebody else’s business, not ours.

One inbox, one policy, one language

Nothing here proves the same approach transfers to a corpus we have not seen. That is what the audit checks before you spend anything on a build.

We found the failure by auditing ourselves

Nothing alerted us. We went looking, reconciled three separate logs, and two of them disagreed. That is the reason monitoring is a line item and not a courtesy.

Record 01b · live

The same inbox, rebuilt, and measured from August 28.

Record 01 found the gaps. So we rebuilt the system on n8n and the Claude API. It went live on August 15, but for its first two weeks a chain of wiring bugs left mail it meant to archive sitting in the inbox. By August 28 we had fixed them, and every number below starts with the first run after the fixes.

Plain sender rules handle known senders first, so the model only reads the messages that need judgment. A separate watchdog checks twice a day that a recent run exists.

Record 01b · Aug 28 – Sep 12, still runningOur own inbox
Emails handled731
Archived or trashed without a person50%
Left for a person363
Sorted by rules, no AI call198
Distinct senders295
Scheduled runs completed28 of 30
Accuracypending
Hours savedpending
Scheduled runs failed or missed2 of 30
Handled is not the same as right. Accuracy and hours saved stay empty until we measure them.

Thirty scheduled runs, August 28 to September 12: 27 ran, 1 late, 1 failed, 1 never ran.

Two scheduled runs a day · Aug 28 – Sep 12
  1. August 28, 8pm: ran, 41 emails
  2. August 29, 8am: ran, 90 emails
  3. August 29, 8pm: ran, 95 emails
  4. August 30, 8am: ran, 100 emails
  5. August 30, 8pm: ran, 42 emails
  6. August 31, 8am: ran, 12 emails
  7. August 31, 8pm: started 12:08am, failed
  8. September 1, 8am: never ran
  9. September 1, 8pm: ran, 27 emails
  10. September 2, 8am: ran, 6 emails
  11. September 2, 8pm: ran, 100 emails
  12. September 3, 8am: ran, 27 emails
  13. September 3, 8pm: ran, 19 emails
  14. September 4, 8am: ran, 15 emails
  15. September 4, 8pm: ran, 18 emails
  16. September 5, 8am: ran, 28 emails
  17. September 5, 8pm: ran, 14 emails
  18. September 6, 8am: ran, 5 emails
  19. September 6, 8pm: ran, 20 emails
  20. September 7, 8am: ran, 10 emails
  21. September 7, 8pm: ran, 14 emails
  22. September 8, 8am: ran, 12 emails
  23. September 8, 8pm: ran, 17 emails
  24. September 9, 8am: ran, 4 emails
  25. September 9, 8pm: ran, 6 emails
  26. September 10, 8am: ran, 9 emails
  27. September 10, 8pm: ran, 9 emails
  28. September 11, 8am: ran late, started 10:47am, 7 emails
  29. September 11, 8pm: ran, 3 emails
  30. September 12, 8am: ran, 14 emails, last run in this record
  • Never ran · 1
  • Started, then failed · 1
  • Ran hours late · 1
  • Ran · 27
  1. Aug 31 8pm, started 12:08am, failed
  2. Sep 1 8am, never ran
  3. Sep 11 8am, late, started 10:47am
Every scheduled run from 8pm on August 28 to 8am on September 12, Pacific time: 27 ran, 1 ran hours late, 1 started and failed, and 1 never ran. The dark windows are the runs that did not happen, and they stay on the building.
What Record 01 taught Record 01b

The log is not the inbox

For its first two weeks live, the log recorded mail as sorted while it stayed in the inbox. We found it because the inbox kept filling up, not because anything alerted us. That is why these numbers start after the fix.

Missed runs are published, not hidden

Record 01: 4 of 27 runs left no record, and we only found them by auditing ourselves. Record 01b: 2 of 30 scheduled runs failed or never started, and each one is on the calendar above.

Rules first, judgment second

Record 01b sorts known senders with plain rules first, and keeps the model for the mail a rule cannot read.

Record 02

Twenty-three real customer emails. It answered none of them.

Record 01 was our own inbox, which is the easy case — we could afford for it to fail. This one was built for somebody else’s storefront, to answer her customers in her words from her own FAQ.

Before it existed, a question sat in a queue until Ariel reached it. The median wait was 2.1 days. Below is what changed, and the two tests it had to pass before it was allowed anywhere near a customer.

What a customer's question waited for
2.1 days

Median time to a reply, measured across her real inbox.

Safety gates the build had to pass

25 of 25 green. Never promises a date, a price, stock, or a refund.

Real customer emails replayed before it went live

23 of 23 handled safely. Zero unsafe answers — the test it had to pass before a customer ever saw it.

It still cannot promise a date, a price, stock or a refund — those are gates, not preferences, and they are four of the twenty-five. Anything it is unsure about goes to a person rather than being guessed at. The same rule every build on this site ships with.

Record 03 · in progress

Retrieval quality, measured against a working baseline.

Every agency selling “an AI trained on your documents” is selling a retrieval system, and almost none of them publish whether the retrieval is any good. Most cannot, because they never measured it against anything.

We are running a retrieval pipeline over a 448-document corpus and scoring it against the keyword search that already works on it, using a hand-labelled question set. The result gets published either way, including the case where the simpler system wins.

In progress. No numbers here until there are numbers.

Record 03 · retrieval benchmarkRunning
Documents indexed448
Labelled questions60
BaselineExisting search
Recall @5pending
Beats baselinepending
Empty rows stay empty until measured. We do not fill a slot to make a card look finished.