Institutions should trust AgriVoice only after consented field speech shows that it can preserve meaning when farmers move between Asante Twi and English, and escalate when it cannot. Clean recordings and read-speech benchmarks cannot prove that a changed word will be caught before it becomes harmful advice.
Consider an illustrative scene in a cocoa-growing community in Ghana. At 6:40 a.m., Kwaku stands beside six tied sacks while a buyer waits near the roadside. He presses record on WhatsApp and asks in Asante Twi, slipping into English for “purchasing season” and then returning to Twi to ask whether he should sell now.
The transcript changes the crucial phrase. AgriVoice receives a question about a different season and selects guidance that does not answer what Kwaku meant.
The buyer may leave. If Kwaku acts on the wrong interpretation, he could make a sale based on guidance that never applied to his question.
Code-switching can change the job of the question
Code-switching is normal speech, not noise around an otherwise tidy sentence. A farmer may use Asante Twi for the situation, English for an institutional term, and Twi again for the decision he needs to make. The switch may happen mid-phrase, with field noise, a short pause or a pronunciation shaped by local usage.
That creates a demanding test for speech recognition. A transcript can look fluent while changing the farmer’s intent. The system might preserve every surrounding word and still lose the one phrase that determines whether the farmer is asking about timing, eligibility, price, spraying or storage.
This matters because AgriVoice does not generate agronomy or purchasing policy from scratch. Its reasoning layer selects from reviewed content blocks. That constraint reduces one source of risk, but it cannot rescue a question whose transcript points to the wrong block.
A plausible sentence is not enough. The transcript must preserve the decision the farmer is trying to make.
Field speech must test failure, not hide it
Neuralis currently uses Meta’s Omnilingual model as its Twi speech-recognition baseline. It beat the project’s fine-tuned Whisper model on read speech, but read speech leaves out much of what matters here: spontaneous phrasing, code-switching, farm vocabulary, background sound and the way someone speaks when another person is waiting for an answer.
Before farmer exposure, the AgriVoice plan requires at least 20 real farmer questions to be recorded and transcribed with consent. Those recordings should include the difficult cases institutions need to see, not only the clearest samples:
- Farmers switching between Asante Twi and English within one request.
- Cocoa terms, institutional language and proper nouns.
- Ambiguous questions that could match more than one reviewed block.
- Recordings with realistic background noise or clipped words.
- Questions outside the approved content set.
- Speech-recognition errors that could change a safe answer into an unsafe one.
Consent must cover more than permission to press record. Farmers need a clear account of why the audio is collected, how it will be used, where it is retained and how deletion works. The deletion process also needs to be tested. A privacy promise that nobody has tried to fulfil remains an untested promise.
The same field recordings should feed the planned direct-Twi versus translation-assisted selection experiment. About 40 questions will be labelled across common, ambiguous, code-switched, out-of-domain and speech-corrupted cases. Reviewers can then compare exact block selection, recall, unsafe extra blocks, escalation accuracy, latency and cost.
A one-question lead in a sample that small does not establish equivalence. Safety errors deserve more weight than a narrow accuracy win.
Trust begins when the system admits uncertainty
Back beside the sacks, Kwaku needs the workflow to detect that “purchasing season” did not survive transcription with enough confidence. The correct turn in this illustrative scene is a pause: AgriVoice declines to select an answer and sends the question to the named extension officer responsible for escalations.
That officer checks the audio rather than relying only on the damaged transcript. Kwaku receives a verified response while the sale is still possible. The six sacks remain his decision to make, now based on the question he actually asked.
This is the same safety principle described in what happens when a cocoa farmer’s question cannot be answered safely and in the comparison of constrained selection, generative answers and human escalation.
Institutions should ask for evidence from the full workflow. Can it recognize enough of the field recording to select the right reviewed content? Does it escalate ambiguous or damaged input? Does a named person receive that escalation? How quickly do they respond? Can the team identify and delete a farmer’s recording when requested?
The two-week pilot is designed to produce those answers with 20 to 50 farmers. Its continue signals include at least 70 percent of questions answered or correctly escalated, zero unsafe pesticide answers, at least 80 percent reported comprehension and an accountable escalation process with a median response below one working day.
If repeated confident errors appear, the workflow must stop or be redesigned. Institutional trust should begin with that willingness to stop.
For Kwaku, the proof is smaller and more concrete. He asks one mixed-language question, and the system either preserves its meaning or puts a person in the loop before six sacks move on the strength of the wrong words.
Comments
No comments yet.