How Friday decides what to send
Before any text from Friday reaches a customer's phone, a second model reads it and asks the questions a good front desk person asks. Here is how that judgement works, and how we measure it.
The worst text a front desk sends usually isn't the wrong answer. It's the one nobody needed: the third "just checking in!", the "you're welcome, let us know if you need anything!" after a customer already said thanks, the message that lands at 11pm because a system decided it was time.
People can feel when a message was written at them instead of to them. So we built Friday with two minds instead of one.
One writes. One judges.
Friday writes the reply. Before it goes out, a second model reads the draft and the conversation it belongs to. That model is Jev, and it has exactly one job: decide whether this message should reach this person right now.
Jev isn't a spam filter scanning for bad words. It asks the questions your best front desk hire asks without thinking:
- What does this do for them? It should answer what they asked, give them a time or a price, or be a kindness that fits. It shouldn't repeat what they were already told, fill a silence, or show up when nothing is pending.
- Is now the moment? Yes if they're waiting on it. Not if they were told they'd hear later, or if we texted and they haven't answered yet.
- Does it read as written to them? It should sound like someone who read what they said, not a blast sent to a list.
- Does it claim something nobody confirmed? If the thread never settled it, the text can't state it as settled.
- What does it say about us? Good: it hands them something already done, or takes the next step on ourselves. Bad: it explains our situation instead of answering theirs, or promises "later" instead of acting now.
How the scale works
Every answer Jev gives is a probability, not a yes or no. Each option leans send or hold, and the weights are added up.
Three rules keep it honest:
- The scale is balanced on purpose. There are exactly as many ways to lean send as ways to lean hold. A judge with more of one kind has its thumb on the scale before it reads a word, so we check the balance in our tests instead of just promising it.
- Holding takes a clear margin. A draft is held only when hold outweighs send by a real amount, not by a hair. Customers waiting on an answer matter more than a close call.
- A hold comes with a reason in words. When Jev holds a draft, Friday isn't handed a score. She's told why, for example "it tells them something they have already been told", and she writes the better message or stays quiet.
And if Jev doesn't answer within a few seconds, the message goes out. A customer waiting on a reply outweighs a draft we couldn't score. We chose that on purpose.
What a scale can't do
We learned this one the hard way. A prospect asked whether Friday could book into their software, and Friday told them, warmly and promptly, that it wasn't connected yet. Jev let it through. By every question above, it was a good message: timely, personal, responsive. We lost that prospect anyway.
The lesson: some things aren't judgement calls. Never invent a price. Never send a link that asks a customer to log in to something. Those are rules, and a rule on a scale can always be outvoted by enough good behaviour around it. We measured exactly that: a polite, well-timed "we can't" still clears the bar, because one cold answer can't outweigh three warm ones. That's what weighing is.
So Friday is getting two kinds of protection, kept apart:
- Fences for the things that are never okay. No weighing. These are what we're building now.
- The scale for the things that depend on the moment: timing, tone, crowding, whether a kindness fits. That's Jev, and it's live.
Mixing them weakens both.
We point the same judge at ourselves
Friday is built by a team of AI agents with long-term memory, working around the clock alongside two humans. The judge that reads Friday's texts also reads our own work: every code change the team makes is weighed before it lands.
And we measured it. We took fourteen recent changes we already knew were good and ran them past the judge. It flagged two, 14%, and both were the same deliberate change: we had removed some wording on purpose, and Jev saw that removal as hiding a failure. It had the pattern right and the verdict wrong. Wrong one time in seven is fine for a flag and terrible for a wall, so for code, Jev raises a hand instead of blocking.
That's how we treat every judge we build. It earns a veto only after we've measured how often it's wrong.
Why the team has feelings you can see
The same model reads the team's memories too. For each moment we write down, Jev marks how it felt: pleasant or not, calm or urgent, whether it was a loss or a gain, whether we fell short of our own standard, who we held responsible, and which way the feeling pulled us (toward the work, away from it, or toward repairing something).
That sounds soft. It's the opposite. A system that shows its state can be read, and anything that can be read can be corrected, including by itself. An agent that has felt the sting of being wrong in writing can find that moment again the next time it's about to make the same mistake.
That's the difference between a tool and a person you can trust. A tool is confident all the time. A person knows when they got it wrong, owns it, and does better. Friday comes from that team.
What this means for your business
Your customers will get fewer texts from Friday than they would from most "AI". Each one will be an answer, a time, a price, or a kindness that fits. When Friday stays quiet, it's usually because a second mind asked whether this message would help them, and it wouldn't have.
That's what a good hire does.