NeuralisNeuralis
← All posts

Kojo’s cocoa question is unclear. Unsafe spray guidance could follow.

A consented farmer voice recording shows how a cocoa question sounds when language switches, a product name arrives in English, and field noise competes with every word. Clean lab samples help establish a baseline, but they cannot show where a real conversation becomes unclear or unsafe.

At 6:40 in the morning, Kojo stands at the edge of his cocoa plot near Kumasi with a phone in one hand and a leaf marked by dark spots in the other. A motorbike passes on the road behind him. He begins in Asante Twi, pauses over the product name on a bottle, says it in English, then asks whether he should spray before rain.

His question contains the detail that matters most: he may be about to act on an answer. If the product name is misheard, if the Twi phrase is mistaken for a different disease symptom, or if the rain question disappears under the motorbike, the wrong response could leave him spraying at the wrong time or using guidance meant for another product. The bottle stays in his hand while the uncertainty remains.

That is why AgriVoice needs consented recordings from farmers before a pilot. The goal is to measure what the system can actually understand, answer safely, or send to a named extension officer. A clean recording cannot carry all of those lessons.

Real questions arrive with language switches and interruptions

A lab sample can be useful: one speaker, a clear prompt, little background sound, and a known transcript. It tells a team whether speech recognition can handle carefully spoken Twi. It does not recreate the way a farmer may talk while walking a plot, checking a pod, speaking to a neighbour, or reading a label.

Cocoa questions can move between Twi and English inside one sentence. Farmers may use an English product name, say “fungicide,” mention COCOBOD, or repeat wording from a label that is hard to pronounce over a voice note. The system must preserve those terms rather than confidently substituting a familiar but wrong word.

Neuralis has already seen why that distinction matters. In an earlier translation check, “kokoo,” meaning cocoa, came back as “chicken.” Fluent output can still be dangerously wrong when a domain noun changes. AgriVoice is designed so that the language model selects from reviewed advice blocks rather than writing agronomy advice from scratch. When the question cannot be matched safely, it should escalate.

But selection can only be as safe as the words it receives. A real recording reveals whether the system captures the product name, the symptom, and the farmer’s actual request together.

Field noise changes the safety decision

The important result from a noisy recording may be an escalation. That is a success when the alternative is a confident answer built on missing information.

Imagine Kojo’s voice note reaches the system with the phrase after the product name clipped by wind. The transcription may still contain enough Twi to suggest a disease question, but it may lose the wording that identifies the bottle or the timing concern. A clean sample would never expose that gap. A field recording makes it visible, measurable, and fixable.

For AgriVoice, the pilot scorecard treats correctly escalated questions as successful answers alongside reviewed responses. The standard is not to answer every message. The standard is to avoid unsafe pesticide guidance and make sure a named extension officer can take over when the system lacks enough certainty.

That boundary matters especially for chemical questions. Three current content blocks remain withheld because dosage, re-entry, and pre-harvest guidance still require agronomist review. The right system behaviour is restraint. The same principle appears in why a pesticide question must reach an extension officer: when the facts needed for a safe answer are absent, the conversation must continue with a person.

A farmer’s recording is personal data. It can contain a voice, phone-linked context, details about a farm, and an urgent question. Consent and deletion behaviour therefore need to be tested before farmer exposure, alongside speech recognition.

The useful record is not an open-ended archive of voices. It is a consented, limited set of questions collected for a stated purpose: assess field speech recognition and the safety of the full voice workflow. The team can compare the recording with a transcription, note where a product name or Twi phrase was lost, and check whether the eventual response was reviewed, understandable, or correctly escalated.

This also gives the pilot partner something concrete to examine. A cocoa-sector partner and their named extension officer can see the kinds of questions entering the queue, the points where audio quality changes meaning, and the cases that need human follow-up. That is a stronger basis for a two-week pilot than a polished demo alone.

The next test is whether Kojo can understand the reply

Kojo’s question has a second half. After the system processes his recording, can he understand the spoken reply well enough to act safely?

Neuralis currently uses an existing Asante Twi VITS voice for the pilot demo. Its measured intelligibility did not meet the long-term acceptance gate, and it has a formal sound shaped by scripture recordings. Field users may find it understandable enough for a limited pilot, or they may show that a conversational recording is urgent. Only real, consented use can answer that honestly.

For Kojo, the best ending may be a short reviewed response that clearly fits his question. It may also be a message that says the system cannot safely confirm the product or spraying conditions and has passed the question to the extension officer. Either way, he should know what happens next before he opens the bottle.

Neuralis

AgriVoice helps Asante-Twi-speaking cocoa farmers ask farming questions by voice and receive answers assembled only from agronomist-reviewed content, with human escalation when the system is unsure.

Try Neuralis

Comments

No comments yet.