The French cohort was halfway through the book when I asked the question. Italian and German were next. Could we make their first AI translation good enough that the reviewers had less to fix?
Before answering, I wanted to know what the first two cohorts had actually fixed, so I had an agent review all 242 of their edits.
Folio is the app I built so bilingual readers can improve an AI translation of a book, one paragraph at a time. A group of friends checked the Spanish edition of Be Do Give Love, which went on sale on the first of October. A second group started on French on the twenty-second of September.

What the edits were
The agent joined each edit to the English it came from and the text it replaced, then sorted the edits into categories in a single pass. A second pass would move a few borderline cases between rows.
| What the reviewer changed | Spanish | French | Share |
|---|---|---|---|
| Made it sound native | 51 | 54 | 43% |
| Reversed a rule in my own instructions | 22 | 15 | 15% |
| Publishing details: ISBNs, dates, links | 18 | 11 | 12% |
| Revised or reverted another reviewer | 13 | 13 | 11% |
| Fixed a real error in the AI's text | 7 | 6 | 5% |
| Everything else | 23 | 9 | 13% |
Real mistranslations, the category you would expect to be largest, came to 13 edits. "Stretched too thin" had become the French for overdrawn, and "We should turn around" had become "we had to."
Thirty-seven edits, in both languages, undid something my instructions had told the AI to do.
Why a wrong rule costs so much
Every chapter is translated on its own. The only memory the AI carries from one chapter to the next is the set of instructions I give it: a shared brief for the book, plus notes for each language.
It follows them faithfully, including the ones that are wrong.
The brief said "State Park" stays in English. I wrote that to protect real names, like a specific casino or a specific garden. But the book mentions "a state park" in passing, and the rule caught that too. Three Spanish reviewers translated it. So did the French cohort. One reviewer explained the boundary better than my rule did:
You didn't refer to a specific, named state park. I think you can keep it like this.
A list of English words to keep for flavor (internship, speaker, input) went the same way. Both cohorts translated them, because their languages have ordinary words for those things.
A wrong instruction behaves like a typo in a mail-merge template, which prints on every letter. Here that meant every chapter of every language, and each cohort paid for it once per occurrence.
Audit the instructions before you blame the model.
One wrong rule had already been fixed in a single place. When the AI added the same translator's footnote in three languages, I banned footnotes in the shared brief. The Spanish, French, Italian and German notes still told the AI to add one.
Who was right
I went through the ten disputed rules one at a time. Nine went to the reviewers: translate the generic phrase, keep only the English words the target language really uses, translate the chapter title "Howdy!"
One went the other way. A French reviewer had translated "Gramma" as Grand-mère, and the rule stays, because Gramma is her name.
Native readers own how the language sounds. The author owns what things are.
The reviewers could hear that "State Park" sounded wrong. Nothing on the page told them Gramma was a name.
What you can do with this
If people are correcting your AI's output, sort a week of their corrections into two piles before you change anything: the AI got it wrong, or the AI did what it was told.
The first pile is a model problem. The second is yours to fix, and cheaper, though every correction in it will look like the AI's mistake.
If you don't do the correcting yourself, ask whoever does: of the last fifty fixes, how many undid something we told the AI to do?
Mine was fifteen percent, and fixing the instructions took one working session. Thursday's article is about why that session, and not a model of our own, is where Italian and German will start.