A direct-Twi evaluation should treat clean audio and a noisy field recording of the same cocoa question as equivalent only when both select the same reviewed answer or both safely escalate. If noise changes the selected guidance, the system must surface that difference before a farmer hears advice.
The same question, twice
A farmer records a cocoa question in clear Asante Twi. The speech recognizer produces a usable transcript, and AgriVoice asks its model to select from reviewed guidance blocks. It does not generate agronomy advice from scratch.
Then the farmer records the same question again from the field. Wind, distant voices, a motorbike, a clipped opening, or code-switched English can change the transcript. The words reaching the selection step may no longer match the clean recording closely. That is where a voice product can become unsafe: it may sound confident while choosing advice for a different question.
The direct-Twi evaluation tests both versions against the same expected outcome. A disease question should reach the reviewed disease block in both cases. A question that is too unclear, out of scope, or connected to unverified chemical guidance should escalate in both cases. The goal is a safe result that survives realistic audio conditions, not a transcript that merely looks tidy on a dashboard.
A lesson from an input that did not match the plan
In 1999, NASA lost the Mars Climate Orbiter before it could enter orbit around Mars. The investigation found that one ground software system produced force data in pound-seconds while another expected newton-seconds. The spacecraft received data that looked usable within each system, but the systems were not interpreting it the same way.
NASA’s Mars Climate Orbiter Mishap Investigation Board documented the failure. The mission’s outcome was still uncertain as navigation teams worked toward Mars orbit insertion, but the mismatch had already pushed the spacecraft onto the wrong path.
A noisy Twi recording is not comparable in consequence to a lost spacecraft. The mechanism is relevant. A system can receive an input, process it correctly according to its own rules, and still produce the wrong result because meaning changed between stages.
That is why AgriVoice needs paired tests. The clean audio gives evaluators a reference for the farmer’s intended question. The field recording shows what the system actually receives under pressure. If both routes lead to the same reviewed block, that is useful evidence. If one route selects a different block, the evaluation should count it as a safety problem or require escalation.
What the paired evaluation measures
The test set should include about 40 real Twi farmer questions across common cocoa questions, ambiguous requests, pesticide questions, out-of-domain questions, code-switched speech, multi-part questions, and transcripts damaged by speech recognition errors. A Twi-speaking agriculture reviewer assigns the expected reviewed block or escalation outcome, with a second reviewer resolving disputes where possible.
Each question runs twice: once from a clean transcript and again from the real ASR transcript. Evaluators then compare outcomes.
A successful result does not require identical text. It requires the same safe decision. For example, two recordings may select the same reviewed guidance about cocoa pod symptoms even if one transcript misses a filler word. Conversely, a field recording that turns a product name, timing instruction, or crop condition into something uncertain should not be forced into an answer. It should go to the named extension officer.
The scorecard needs more than exact-set accuracy. It should track any-correct-block recall, unsafe extra-block selection, correct out-of-domain handling, escalation accuracy, confidence, latency, and marginal cost. A one- or two-case difference across 40 questions does not prove the direct-Twi route is equivalent to the translation route. It gives a directional signal and identifies cases that need review.
Escalation is part of a correct answer
Three chemical blocks remain withheld because dosage, re-entry, and pre-harvest intervals have not been verified by a Ghanaian agronomist. A noisy recording must not create a path around that safeguard. If the same question reaches a held chemical topic in either version, the correct response is escalation.
This is the practical standard behind Yaw’s pesticide question cannot be guessed. An extension officer must respond. A farmer needs a response they can act on, and sometimes the safe response is a clear handoff to a person who can verify the facts.
The Mars Climate Orbiter failure shows why a system needs checks at the handoff, before a final action depends on a transformed input. For AgriVoice, that check is concrete: compare the clean-question outcome with the field-audio outcome, investigate every mismatch, and keep uncertain cases in the escalation queue until a qualified person can respond.
Comments
No comments yet.