Everything here is a number we produced ourselves.
No client logos we have not earned. No testimonials we did not receive. No industry statistics borrowed from a report. When we have not measured something, we say so and leave the space empty.
This is the shortest page on most agency websites. We would rather it be the longest on ours.
Autonomous email triage, fourteen days on a live inbox.
A scheduled AI agent ran twice a day against a real working inbox for two weeks. It read the body of every message, not just the sender, then filed, flagged, archived, or escalated according to a written policy it also helped build.
It caught things a filter cannot. A phishing reply dressed as a support thread, routing identity verification to an off-brand domain. Two marketing blasts disguised as personal email, one of which was only detectable from a tracking beacon in the body. A single vendor family turned out to use seven sender addresses that needed routing three different ways, which is why sender rules were the wrong abstraction and judgment was the product.
It also had a real gap. 4 of 27 scheduled runs left no record. We checked two: both had genuinely failed — a timeout, a connection error. The other two, we still don’t know. What was missing wasn’t the error itself — it was the alert. On a personal inbox that is a shrug. On a client’s lead notifications it is a lost deal, which is why every build we ship now alerts on absence.
This is the system that produced those numbers — including the one that failed.
Below is the actual shape of it. The first state is what we built. The second is what we run now, and the difference between them is the only reason the gap above is on this page at all.
and nothing said so
A run can fail before it ever writes a usable error. Every monitor we had was watching the output, so when a run didn’t complete, nothing had an opinion about it. The fix is the node at the bottom, and the direction of its arrow: it watches the trigger, not the result. That is now a build standard on every system we ship.
Measured on our own operations before we sold anything to anyone. The gap is on this card on purpose. Two of the four we checked were real errors nobody was alerted to; the other two, we still don’t know. An agency that has never found a failure has never looked.
A number without its limits is marketing.
This was our inbox, not a client’s
Around six threads a day. A shared support inbox running two hundred a day is roughly thirty times the volume against the same build. We publish the multiplier rather than a promise, and your audit produces your real number.
Fourteen days is a trial, not a track record
Long enough to find a silent failure mode. Not long enough to claim reliability. Record 02 is the harder test: somebody else’s business, not ours.
One inbox, one policy, one language
Nothing here proves the same approach transfers to a corpus we have not seen. That is what the audit checks before you spend anything on a build.
We found the failure by auditing ourselves
Nothing alerted us. We went looking, reconciled three separate logs, and two of them disagreed. That is the reason monitoring is a line item and not a courtesy.
The same inbox, rebuilt, and measured from August 28.
Record 01 found the gaps. So we rebuilt the system on n8n and the Claude API. It went live on August 15, but for its first two weeks a chain of wiring bugs left mail it meant to archive sitting in the inbox. By August 28 we had fixed them, and every number below starts with the first run after the fixes.
Plain sender rules handle known senders first, so the model only reads the messages that need judgment. A separate watchdog checks twice a day that a recent run exists.
Thirty scheduled runs, August 28 to September 12: 27 ran, 1 late, 1 failed, 1 never ran.
Two scheduled runs a day · Aug 28 – Sep 12- August 28, 8pm: ran, 41 emails
- August 29, 8am: ran, 90 emails
- August 29, 8pm: ran, 95 emails
- August 30, 8am: ran, 100 emails
- August 30, 8pm: ran, 42 emails
- August 31, 8am: ran, 12 emails
- August 31, 8pm: started 12:08am, failed
- September 1, 8am: never ran
- September 1, 8pm: ran, 27 emails
- September 2, 8am: ran, 6 emails
- September 2, 8pm: ran, 100 emails
- September 3, 8am: ran, 27 emails
- September 3, 8pm: ran, 19 emails
- September 4, 8am: ran, 15 emails
- September 4, 8pm: ran, 18 emails
- September 5, 8am: ran, 28 emails
- September 5, 8pm: ran, 14 emails
- September 6, 8am: ran, 5 emails
- September 6, 8pm: ran, 20 emails
- September 7, 8am: ran, 10 emails
- September 7, 8pm: ran, 14 emails
- September 8, 8am: ran, 12 emails
- September 8, 8pm: ran, 17 emails
- September 9, 8am: ran, 4 emails
- September 9, 8pm: ran, 6 emails
- September 10, 8am: ran, 9 emails
- September 10, 8pm: ran, 9 emails
- September 11, 8am: ran late, started 10:47am, 7 emails
- September 11, 8pm: ran, 3 emails
- September 12, 8am: ran, 14 emails, last run in this record
- Never ran · 1
- Started, then failed · 1
- Ran hours late · 1
- Ran · 27
- Aug 31 8pm, started 12:08am, failed
- Sep 1 8am, never ran
- Sep 11 8am, late, started 10:47am
The log is not the inbox
For its first two weeks live, the log recorded mail as sorted while it stayed in the inbox. We found it because the inbox kept filling up, not because anything alerted us. That is why these numbers start after the fix.
Missed runs are published, not hidden
Record 01: 4 of 27 runs left no record, and we only found them by auditing ourselves. Record 01b: 2 of 30 scheduled runs failed or never started, and each one is on the calendar above.
Rules first, judgment second
Record 01b sorts known senders with plain rules first, and keeps the model for the mail a rule cannot read.
Twenty-three real customer emails. It answered none of them.
Record 01 was our own inbox, which is the easy case — we could afford for it to fail. This one was built for somebody else’s storefront, to answer her customers in her words from her own FAQ.
Before it existed, a question sat in a queue until Ariel reached it. The median wait was 2.1 days. Below is what changed, and the two tests it had to pass before it was allowed anywhere near a customer.
Median time to a reply, measured across her real inbox.
25 of 25 green. Never promises a date, a price, stock, or a refund.
23 of 23 handled safely. Zero unsafe answers — the test it had to pass before a customer ever saw it.
It still cannot promise a date, a price, stock or a refund — those are gates, not preferences, and they are four of the twenty-five. Anything it is unsure about goes to a person rather than being guessed at. The same rule every build on this site ships with.
Retrieval quality, measured against a working baseline.
Every agency selling “an AI trained on your documents” is selling a retrieval system, and almost none of them publish whether the retrieval is any good. Most cannot, because they never measured it against anything.
We are running a retrieval pipeline over a 448-document corpus and scoring it against the keyword search that already works on it, using a hand-labelled question set. The result gets published either way, including the case where the simpler system wins.
In progress. No numbers here until there are numbers.