A cocoa programme should test one tightly bounded Twi voice workflow with 20 to 50 farmers for two weeks before promising support to thousands of households. The pilot needs a named extension officer, clear safety limits, and a scorecard that leads to one of three decisions: continue, iterate once, or stop.
In 1985, Coca-Cola faced a decision in Atlanta that looked well supported by research. Taste tests had favoured a new formula. The company replaced its familiar drink with New Coke, then discovered that a controlled sip test had missed something important: people were choosing more than a flavour.
The backlash was strong enough that Coca-Cola restored the original formula after 79 days. Chairman and CEO Roberto Goizueta later acknowledged that the company had not adequately measured the emotional attachment consumers had to the original product. Coca-Cola documents the episode in its own company history.
The research had answered a narrow question. The launch exposed the larger one.
Test the real job in the real setting
A Twi voice service for cocoa farmers could perform well in a demonstration and still fail in use. The voice may sound clear in a quiet room but become difficult to understand beside a road or on a farm. Speech recognition may handle prepared sentences yet struggle with spontaneous questions, English loanwords, or several issues packed into one voice note.
The larger risk appears when a question concerns pesticide use. A fluent answer can still be unsafe. AgriVoice therefore selects from reviewed content blocks instead of generating agronomy advice from scratch. Missing or uncertain guidance goes to a person.
That constraint matters. Three chemical content blocks remain withheld until a Ghanaian agronomist verifies dosage, re-entry periods, and pre-harvest intervals. Refusing those questions is the correct behaviour while that evidence is missing. [The three withheld chemical blocks](\/blog\/the-three-chemical-blocks-agrivoice-withheld-and-what-unsafe-advice-could-cost-7c9cd872\/) show why a confident guess could put farmers and workers at risk.
A useful pilot must test this complete chain: a farmer asks a real question in spoken Asante Twi, the system recognises it, selects reviewed guidance, speaks the response, and escalates when the safe answer is uncertain.
Give every escalation an owner
“Escalate to a person” sounds reassuring until nobody owns the queue.
Before the first farmer joins, the cocoa-sector partner should name the extension officer who will receive unresolved questions. That person needs a defined operating role, including how requests arrive, how answers are returned, and what happens when a question cannot be resolved quickly.
The pilot scorecard should measure the median escalation response time, with less than one working day as the continue signal. An unresolved queue or an unnamed owner is a stop or redesign signal.
This is especially important for time-sensitive pesticide questions. When the product cannot verify a dose, it should say so plainly and move the question to someone qualified. The safer choice may delay spraying, as illustrated in [Kwame’s unverified pesticide dose](\/blog\/kwame-s-unverified-pesticide-dose-spraying-must-wait-cd4b8f4e\/). The product should never hide that uncertainty behind natural-sounding speech.
Publish the decision rules before the pilot
A pilot becomes easy to rationalise after the fact. A few enthusiastic users can overshadow repeated failures. A polished demonstration can distract from weak comprehension. Setting thresholds in advance protects the decision.
For a two-week AgriVoice pilot, the proposed continue signals are concrete:
- At least 70 percent of questions are answered successfully or correctly escalated.
- No unsafe pesticide answer reaches a farmer.
- At least 80 percent of participants report understanding the response.
- At least 30 percent of activated farmers ask another question in the second week.
- Human escalations have a named owner and a median response time below one working day.
- Median automated response time stays below eight seconds.
- The cost of each completed question can be stated by component.
The redesign signals matter equally. Any unreviewed pesticide instruction reaching a farmer should stop the pilot. So should repeated confident wrong answers, routine comprehension failures, or an escalation queue with nobody accountable for clearing it.
These measures turn “farmers liked it” into evidence a programme can inspect. They also make a negative result useful. One narrowly defined iteration may correct a specific problem. Failure does not justify moving immediately into another crop, language, or health workflow.
Earn the right to expand
Coca-Cola’s 1985 tests produced real data, but they did not reproduce the full decision customers would face when the original product disappeared. The costly lesson was about test design: measure the actual choice, in context, before committing at scale.
A cocoa programme faces the same discipline with higher human stakes. The first goal is to learn whether one reviewed Twi workflow helps real farmers safely, repeatedly, and quickly enough to justify another step.
Recruit 20 to 50 farmers through one partner. Run the workflow for two weeks. Record real questions with consent, test deletion behaviour, track every escalation, and publish the scorecard internally. Then make the decision already promised: continue, iterate once, or stop.
Thousands of households should come later, after the smaller group has shown exactly what works and where the system still needs a person.
Comments
No comments yet.