If the system cannot greet a customer properly, the rest of the demo does not matter.
Use the same words
We gave ElevenLabs, Google, OpenAI, and Addis AI the same Amharic customer support script. No translated summaries. No different prompts chosen to flatter each system. The input stayed fixed so the output could be heard side by side.
Three systems struggled with the greeting. For an Amharic speaker, that failure arrives immediately. You do not need a chart to hear it.
A customer support greeting is a serious test
Customer support speech has a job. It has to pronounce the words correctly, carry the right tone, and sound clear enough that a caller trusts the next sentence.
A voice that works for an isolated phrase can still fail on pacing, emphasis, names, or sentence-level rhythm. A full support script puts those problems in the open.
Pronunciation
Are the Amharic words spoken correctly?
Rhythm
Does the sentence move like natural speech?
Tone
Would the delivery work in a customer interaction?
Clarity
Can a listener follow it without compensating for the model?
What this comparison does and does not prove
This was not a broad benchmark of every voice, language, and setting available from each provider. It was one controlled product example with one script.
That narrowness is useful. It answers a practical question: could we put this voice in front of an Amharic-speaking customer for this task? Evaluation has to include that question, even when a larger benchmark also exists.
Listen to the original comparison
The evidence is the audio. The original post contains the side-by-side comparison in the order it was tested.