Field-speech evaluation must capture how farmers naturally begin speaking, including clipped openings, background noise, code-switching and missing context. Rehearsed prompts can make a voice system look accurate while hiding the failures that appear in the first five seconds of a real farm conversation.
A farmer reaches the gate beside a working cocoa farm, presses record and begins halfway through the thought:
“...na ɛno na ɛhyɛ aseɛ bio.”
No greeting. No neat description of the crop. The word that identifies the disease may have appeared before recording started, or the farmer may assume the listener already knows what “it” means. A motorbike passes. Someone nearby answers in English. The farmer continues in Asante Twi, pauses, then names a product using the English word printed on its container.
That opening is inconvenient for a test script. It is also exactly what the system must understand.
The equipment had to fit the problem in front of them
In April 1970, the Apollo 13 crew faced rising carbon dioxide inside the lunar module after an oxygen tank explosion forced them to abandon the planned Moon landing. The command module carried square lithium hydroxide canisters. The lunar module needed round ones. The available equipment could remove the carbon dioxide, but the parts did not fit.
NASA engineer Ed Smylie and a team in Houston worked out how the astronauts could connect the incompatible components using materials already aboard the spacecraft. NASA’s account documents the improvised adapter and the uncertainty surrounding the crew’s safe return.
The useful lesson is smaller than Apollo 13’s stakes but similar in shape: capability under ideal conditions means little when the input does not fit. The engineers had to solve the problem using the hardware actually inside the spacecraft. A voice system must be evaluated using the speech farmers actually produce.
A clean studio recording asks, “What should I do if I see signs of black pod disease on my cocoa farm?”
A field opening may sound closer to, “It has come again near the lower side,” followed later by the words that identify the affected pods. The system receives fragments, local references and noise. Recognition quality depends on handling that sequence, not merely transcribing a complete sentence read from a sheet.
Rehearsed prompts hide the hardest failures
Read-speech tests remain useful. They reveal whether a model can recognize known words under controlled conditions, and they make comparisons repeatable. Neuralis used held-out sentences to measure its speech models and found that Meta’s Omnilingual ASR outperformed its fine-tuned Whisper baseline on read speech.
That result establishes a baseline. It does not establish farm readiness.
Natural openings introduce different failure modes. A speaker may begin before the recorder is ready. The first noun may disappear under wind, tools or another voice. Pronouns may refer to something visible to the farmer but invisible to the system. English words may appear inside Twi because that is how the farmer knows a pesticide, disease or institution.
Those details can change the selected answer. They can also change whether the system should answer at all.
AgriVoice is designed to select reviewed content blocks rather than compose agronomy advice. If the question remains ambiguous, or if the relevant guidance has not been cleared, it should escalate to a person. That safety boundary matters most when speech arrives in the untidy form that encourages a model to fill gaps. The same principle explains why unverified pesticide guidance must remain withheld.
Capture the first five seconds without coaching them away
A useful field-speech set should preserve the opening exactly as it happened. Do not restart every recording because the farmer hesitated, began mid-sentence or mixed languages. Those moments belong in the evaluation data.
For the AgriVoice pilot, at least 20 real farmer questions need to be recorded and transcribed before exposure to the wider cohort. The set should include natural variation: short questions, long explanations, ambiguous references, pesticide terms, code-switching and speech affected by the recording environment. Consent and deletion behavior must be tested alongside recognition.
Score more than word error rate. Record whether the correct reviewed block was selected, whether an unsafe extra block appeared, whether an uncertain question escalated and how long the complete response took. Run the selection comparison again on real ASR transcripts, because a method that performs well on typed Twi may fail after the recognizer alters a critical word.
The first five seconds deserve their own review. Mark whether recording began late, whether the subject appeared only once and whether background sound covered the words carrying the question’s meaning. Those observations tell the team what to repair: recording prompts, interface timing, recognition or escalation logic.
Test the materials already on board
Apollo 13’s carbon dioxide problem could not be solved by assuming the astronauts had different canisters. Houston’s procedure had to work with the materials already available.
Field evaluation requires the same discipline. Farmers should not have to learn a test writer’s sentence structure before they can ask for help. The workflow must cope with their real openings, or recognize when it lacks enough information and hand the question to the named extension officer.
The next recording session should begin with one simple instruction: ask the question as you would normally ask it. Keep the partial starts. Keep the pauses. Then inspect the first five seconds before anyone cleans the transcript.
Comments
No comments yet.