Last updated July 13, 2026
How accurate is Zonlo's feedback?
Feedback you can't trust is worse than no feedback. This page explains how Zonlo checks what you say, how we test that checking, what the numbers mean, and what we still haven't proven. It is a look under the hood, not a sales page.
What Zonlo is trying to do
Zonlo is a private space to rehearse speaking out loud before a real conversation happens. It is built for people who know the words but freeze in the moment, so everything follows from two rules: the feedback has to earn your trust, and it can never punish a right answer. Telling someone who already freezes when they speak that a correct reply was wrong would recreate the exact harm this app exists to remove.
That's why this page exists. If we're asking you to speak into an app at your most self conscious, you deserve to know what happens to your voice, what judges your reply, and how often that judge is right.
How the checking works
Every reply you say goes through the same five steps, and the whole thing is built so the worst mistake, failing an answer that was actually fine, is the one we guard against hardest.
- Your iPhone turns your voice into text by itself. The audio never leaves your phone and is never saved anywhere. Only your words, as text, go any further.
- That text travels to our server over an encrypted connection, along with which conversation you're in and which reply you're on. Nothing else about you is needed to check a sentence.
- Japanese replies are pulled apart by a grammar tool. Kagome, an open tool anyone can inspect, splits the sentence into words and works out the readings, the furigana and romaji you see in your feedback. That part is fixed rules, not AI guesswork.
- An AI marks the reply against what a good answer looks like in that moment: did what you said work, is there a more natural way to put it, and did you sound as polite as the moment needed. We tried feeding the grammar tool's answer into the politeness judgment and it made the AI too jumpy, flagging replies that were casual but perfectly fine, so for now the AI makes that call on its own. Making it better is live work.
- The result gets checked before you see it. If the AI fails, takes too long, or hands back something we can't read, Zonlo tells you plainly that it couldn't check this one rather than guessing or showing you a wrong red mark. You are never left waiting.
What each number means
- Cases we test against
- How many real learner answers we run the checking over. Each one is something a person actually said, with the right verdict decided by a human, and they come in clusters: several ways of saying the same thing that are all fine, so we can catch a strict marker wrongly knocking one of them.
- Pass and correctness
- Across all those cases, how often the app's pass or fail call, and its call on whether the answer was right, match what the person decided. This is the headline number: does the app agree with a careful human.
- Politeness judgments
- How often the app gets the politeness level right, for example spotting casual Japanese where the moment called for です/ます. This is our weakest number and the one we're working on hardest. The AI makes this call itself, and its misses split between being too harsh on short polite replies and too easy on blunt ones.
- Correct replies we wrongly failed
- Answers that were fine but got marked as wrong. We track this on its own because it is the most damaging mistake Zonlo can make. Telling someone who already freezes when they speak that a good reply was wrong is exactly the harm this app exists to prevent. Our target is zero and we are not there: this run has 6, mostly on the newer later replies in a conversation, which we are still tuning. Closing them is the top thing we're working on.
How we test it
Every change to the checking, whether it's the wording we give the AI, the AI itself, or the code around it, runs over the whole test set before it ships. If it does worse than what's live, it doesn't ship. That's the whole rule, and it has already blocked upgrades that looked better on paper.
The test set leans on the hard cases on purpose: answers that are right but worded differently, politeness calls sitting right on the line, and replies that work even though they're unusual. Anything looks good on easy cases. These are the ones that decide whether the feedback is worth trusting.
An outside check on politeness
There's a trap in testing an AI against answers you wrote yourself. If it agrees with you, all you've learned is that you agree with yourself, not that either of you is right. Politeness is where this bites hardest, because it's the call we're least sure of and, in Japanese, the one we can't yet fully check with a native speaker.
So for politeness we brought in a second opinion that came from neither us nor any AI: three public collections of sentences where the politeness or formality was labeled by researchers, professional translators, or paid raters. We wrote a small checker for each language using grammar rules alone, no AI, and scored how often it agrees with those human labels:
- Spanish, 96% agreement, against Amazon's formality dataset, where the formal and informal versions were written by professional translators. Tú versus usted is a grammar switch, so rules we wrote by hand come very close to the humans.
- Japanese, 79% agreement, against the KeiCO collection, where researchers labeled the honorific level. Keigo uses a fixed set of forms, and the misses bunch up on the rarest, most deferential ones.
- English, 51% agreement, against the Pavlick and Tetreault formality dataset. English formality is a slider, not a switch, so a rules checker only gets about halfway. That is exactly why, for English, we lean on an outside formality model trained on human judgments instead of these rules.
These are agreement rates between our rules checker and outside human labels. They are not the app's own accuracy. The point is that our weakest number now has an anchor that didn't come from us or from an AI.
Current results
Measured on Gemini 2.5 Flash Lite, last run July 13, 2026. Percentages drawn from a test set this size are a direction, not a guarantee. One new hard case can move them, which is exactly why we print the number of cases right next to them. These are our own measurements against answers we labeled ourselves. Nobody has an industry standard test for this kind of feedback, so we can't compare Zonlo to other apps here, and about half the set, the Japanese half, hasn't been read over by a native speaker yet. Take these as our honest internal number, not a proven or comparable claim.
What we still haven't proven
- The test set is still growing. Five hundred cases is enough to catch things breaking and to keep us honest. It is not enough to promise a decimal point. It grew from two hundred, and adding harder cases pulled the headline numbers down a little, which is exactly what an honest test set is supposed to do. It will keep growing.
- The Japanese answers are still being checked against native sources. The correctness and politeness labels a person gave the Japanese cases are being checked against native material before we treat them as settled.
- Later replies are new to the test set. Zonlo's conversations keep going past the first line, and the replies that come after it have only recently been covered. Expect this part to firm up as those cases grow.
- Hearing you has its own error rate. The numbers above cover the marking, not the listening. Your phone occasionally mishears a word, and one wrong word can change a verdict. We kept the listening on your phone anyway, because your voice staying there is not up for debate.
Why we publish this
Language apps love the word "AI" and hate showing their work. Zonlo is built for people who freeze when they speak, and that only works if the feedback earns your trust. So we publish how many cases we test against, what we miss, and where the gaps are, and we update this page as they change. The date at the top is the date of the numbers, not of the writing. For plain language detail on how AI is used and where it falls short, read our AI Disclosure.
Curious how this compares to other apps? Read Zonlo vs Duolingo and Zonlo vs Speak, or see how a conversation works.
Rehearse it here first.
Zonlo is on the App Store. Free to download, and your first five conversations in each language are free.