Field records · measured, not projected

Everything here is a number we produced ourselves.

No client logos we have not earned. No testimonials we did not receive. No industry statistics borrowed from a report. When we have not measured something, we say so and leave the space empty.

This is the shortest page on most agency websites. We would rather it be the longest on ours.

Record 01 · complete

Autonomous email triage, fourteen days on a live inbox.

A scheduled AI agent ran twice a day against a real working inbox for two weeks. It read the body of every message, not just the sender, then filed, flagged, archived, or escalated according to a written policy it also helped build.

It caught things a filter cannot. A phishing reply dressed as a support thread, routing identity verification to an off-brand domain. Two marketing blasts disguised as personal email, one of which was only detectable from a tracking beacon in the body. A single vendor family turned out to use seven sender addresses that needed routing three different ways, which is why sender rules were the wrong abstraction and judgment was the product.

It also failed. Roughly 15% of scheduled runs died silently. Not errors: absences. The agent never reached the mailbox and nothing anywhere reported a problem. On a personal inbox that is a shrug. On a client’s lead notifications it is a lost deal, which is why every build we ship now alerts on absence.

Record 01 · 14-day live trialOur own inbox
Emails dispositioned146
Senders learned automatically54
Phishing attempts flagged1
Disguised bulk mail caught2
False archives of important mail0
Scheduled runs27
AI cost, 14 days$0.79
Cost per email classified~0.2¢
Runs that failed silently15%
Measured on our own operations before we sold anything to anyone.The 15% is on this card on purpose. An agency that has never found a failure has never looked for one.
The honest part · what this record does not show

A number without its limits is marketing.

This was our inbox, not a client’s

Around six threads a day. A shared support inbox running two hundred a day is roughly thirty times the volume against the same build. We publish the multiplier rather than a promise, and your audit produces your real number.

Fourteen days is a trial, not a track record

Long enough to find a silent failure mode. Not long enough to claim reliability. Record 02 runs longer by design.

One inbox, one policy, one language

Nothing here proves the same approach transfers to a corpus we have not seen. That is what the audit checks before you spend anything on a build.

We found the failure by auditing ourselves

Nothing alerted us. We went looking, reconciled three separate logs, and two of them disagreed. That is the reason monitoring is a line item and not a courtesy.

Record 02 · in progress

Retrieval quality, measured against a working baseline.

Every agency selling “an AI trained on your documents” is selling a retrieval system, and almost none of them publish whether the retrieval is any good. Most cannot, because they never measured it against anything.

We are running a retrieval pipeline over a 448-document corpus and scoring it against the keyword search that already works on it, using a hand-labelled question set. The result gets published either way, including the case where the simpler system wins.

In progress. No numbers here until there are numbers.

Record 02 · retrieval benchmarkRunning
Documents indexed448
Labelled questions60
BaselineExisting search
Recall @5pending
Beats baselinepending
Empty rows stay empty until measured. We do not fill a slot to make a card look finished.