A buying shed can become a speech lab when farmers record real questions there with informed consent. Those recordings reveal accents, code-switching, background noise, interruptions and safety ambiguity that a studio dataset cannot reproduce.
One farmer raises a phone and records the question already moving through the shed. Other voices carry behind him. Cocoa terms sit beside English product names. A number may be spoken in Ghanaian English, while the rest arrives in Asante Twi. The recording is short, but it contains more useful evidence than another hour of carefully read studio prompts.
The clue John Snow could not find in a laboratory
In 1854, physician John Snow faced a problem in Soho, London. Cholera was killing residents, but the dominant explanation blamed contaminated air. Snow suspected water. Suspicion alone could not identify the source or support action while people were still dying.
He mapped deaths around the Broad Street water pump and investigated where affected households obtained their water. The local pattern made the risk visible. Authorities removed the pump handle, and the episode became a landmark in epidemiology.
Snow documented the investigation in the second edition of On the Mode of Communication of Cholera, published in 1855. His evidence came from the conditions in which people actually lived: their streets, homes and water choices. A controlled setting could have examined water samples, but it could not have shown the same human pattern.
The buying shed serves a similar role for speech technology. The model may perform well on sentences read into a quiet microphone. The field recording reveals whether it can understand the person who needs an answer, using the words that person naturally chooses, in the place where the question arises.
Real questions expose data-shaped failures
Neuralis currently uses Meta’s Omnilingual model as its Twi speech-recognition baseline. It beat the team’s fine-tuned Whisper model by 59 percent on read speech. That result supports a direction, but read speech does not settle whether AgriVoice can understand cocoa farmers in use.
A studio speaker knows a recording is coming. They can face the microphone, repeat a sentence and pronounce every word carefully. A farmer at a buying shed may begin halfway through a thought because everyone nearby already knows the problem. They may mention `kokoo`, switch to an English product name, state a dosage as a number and refer to a symptom using a local expression absent from the test set.
The differences matter. Machine translation once turned `kokoo`, cocoa, into “chicken.” A fluent output can still carry the wrong subject. When the system provides agricultural guidance through speech, that type of error cannot be treated as a harmless typo.
Real recordings show where failures come from. The speech recognizer may lose a word. The transcript may be accurate while the content selector chooses the wrong reviewed block. The question may omit the product, timing or field condition needed for a safe answer. Those are different failures, and each needs a different response.
Consent turns recordings into usable evidence
Field speech must be collected deliberately. For the AgriVoice pilot, the entry gate calls for at least 20 real farmer questions to be recorded and transcribed before exposure. Consent and deletion behaviour must also be tested.
That means explaining what will be recorded, why it is needed and how someone can request deletion. A cocoa-sector partner should own recruitment, while a named extension officer receives questions the system cannot answer safely. Audio retention and phone-number storage also need review under Ghana’s Data Protection Act before commercial use.
The recordings should cover more than easy questions. A useful field set includes common agronomy questions, ambiguous requests, code-switching, pesticide questions, out-of-domain requests and speech corrupted by the actual recording environment. The transcript should preserve what was said rather than quietly rewriting it into cleaner Twi.
This evidence also protects farmers when uncertainty matters most. AgriVoice selects from reviewed content instead of inventing agronomy. Chemical guidance remains withheld until a Ghanaian agronomist verifies dosage, re-entry and pre-harvest intervals. When the system cannot find a safe match, the right outcome is escalation, as explained in What Should AgriVoice Do When It Cannot Safely Match a Pesticide Question?.
Bring the evaluation to the shed
The first practical step is small: collect 20 to 30 consented questions from real farmers, transcribe them faithfully and score the existing recognizer against them. Do not fine-tune yet. First identify whether the errors cluster around noise, accents, numbers, code-switching or cocoa vocabulary.
Then run those transcripts through the full path. Measure whether AgriVoice answers from reviewed content, escalates correctly, avoids unsafe extra blocks and responds quickly enough to preserve a spoken exchange. Repeat the evaluation with the raw speech-recognition output, because a selector that succeeds on perfect transcripts may fail on the words the recognizer actually produces.
John Snow’s map mattered because it connected evidence to the place where exposure occurred. Neuralis needs the same discipline. The buying shed supplies the conditions, consented recordings supply the evidence, and the pilot scorecard determines whether AgriVoice should continue, receive one focused iteration or stop.
The next recording should capture the question exactly as the farmer asks it, background voices and all.
Comments
No comments yet.