Home/Tests/Japanese homophones

Native Japanese context test

Can AI Transcribe Japanese Homophones?

Four transcription tools heard the same sentence using three versions of hashi. Only one selected the correct kanji for all three meanings.

Recorded in JapanSame audio for every toolChecked by a native speaker
Independent hands-on test

We wrote and recorded the Japanese source ourselves, used the same audio, and compared the important words with the written script. Product links on this page are direct links; we do not earn a commission from them yet.

One sound, three different meanings

橋の端で箸を落とした。

The natural English meaning is “I dropped my chopsticks at the edge of the bridge.” A transcription system cannot solve this by matching sound to a single spelling. It must understand the surrounding context.

hashi

Edge. The particle で marks where the action happened.

hashi

Chopsticks. The particle を marks the object that was dropped.

What each AI transcription tool wrote

We scored only the three target kanji in the phrase. A correct result had to preserve bridge, edge, and chopsticks in their intended positions.

ToolObserved outputScoreNative review
Notta橋の端で箸3/3All three meanings correct
Otter.ai端の端で箸2/3The first hashi was wrong
Sonix橋の橋で橋1/3Two contextual kanji were wrong
TurboScribe橋の橋で橋1/3Two contextual kanji were wrong

WHAT WORKED

  • Notta: It used bridge, edge, and chopsticks in the intended positions.
  • Otter.ai: It wrote edge where the sentence required bridge.

WHAT FAILED

  • Sonix: Only the first hashi—bridge—matched the intended meaning.
  • TurboScribe: Its clean-looking transcript still changed edge and chopsticks to bridge.

Detailed product findings: Notta review, Sonix review, and TurboScribe review.

A polished Japanese transcript can still be wrong

The Sonix and TurboScribe outputs looked like normal Japanese at a glance. That is exactly why contextual errors are risky for someone who cannot read Japanese: the wrong sentence may not look broken.

INTENDED MEANING

bridge → edge → chopsticks

Sonix and TurboScribe effectively produced “bridge → bridge → bridge.” The text remained readable, but two of the three real-world meanings were lost.

Homophones are only one failure mode. Our broader tests also found errors in prices, quantities, destinations, speaker labels, and mixed letter-number codes. Use this test as a warning sign—not as the only basis for choosing a tool.

What this result does—and does not—prove

  1. 1
    Same source audio.

    Every completed service received the same original recording made in Japan.

  2. 2
    Native contextual check.

    We compared the three target words with the intended written script.

  3. 3
    One controlled sample.

    Different voices, microphones, sentence context, or product updates may produce different results.

  4. 4
    No universal winner from one sentence.

    Our overall comparison also uses five other recordings and multiple failure types.

SEE THE BROADER EVIDENCE

Compare accuracy across six real Japanese recordings.

Numbers, names, corrections, speakers, noise, and business codes included.

See the full comparison →

Japanese homophone transcription FAQ

Can AI transcription understand Japanese homophones?

Sometimes. In this controlled recording, Notta selected all three intended kanji for hashi. Otter.ai selected two, while Sonix and TurboScribe selected one. This is one test sentence, not a universal accuracy ranking.

Why is hashi difficult for speech-to-text tools?

Japanese has many words that share a pronunciation. 橋, 端, and 箸 can all be read as hashi, so the system must use the surrounding grammar and meaning to choose the correct kanji.

Which tool was best in this Japanese homophone test?

Notta was the only tool to score 3/3 in this specific sentence. Our broader recommendation also considers numbers, names, corrections, multiple speakers, noise, and business identifiers.