Independent testing from Japan

Find the AI tool that actually understands Japanese.

We test transcription tools with real Japanese audio, then a native Japanese reviewer checks every number, name, and kanji choice by hand.

Updated September 5, 2026 · Eight controlled recordings

LAB TEST 001–006 EARLY RESULTS

CONTEXT TEST

橋の端で箸を落とした。

Clips8
Tools5
Leader errors4
LeaderNotta

Preliminary—not a final winner. More real-world tests are coming.

01 Recorded in Japan

02 Reviewed by a native speaker

03 Same audio for every tool

04 Weaknesses published

Early benchmark results

Notta leads after eight controlled recordings.

These are hands-on results from eight short controlled recordings. The new workplace test exposed a shared 制作/製作 error even with strong context.

What counts as a major error?A mistake that changes meaning or a numeric value—not a harmless formatting choice.
NEW MARKET GUIDEJapan AI speech-to-text tool market: prices, apps, APIs, and buyer fit
See Japan's speech-to-text market map →
NEW MEETING GUIDEBest Japanese meeting transcription tools, tested with two speakers and noise
Choose a meeting tool →
NEW BUYER'S GUIDEJapanese transcription services: AI vs human, prices, accuracy, and use cases
Choose the right service →
NEW WORKPLACE TESTCan AI tell creative production from physical manufacturing in Japanese?
See the three-tool result →
NEW HOMOPHONE SERIESCan AI choose the right kanji for eight everyday homophones?
See the three-tool result →
NEW CONTEXT TESTCan AI tell bridge, edge, and chopsticks apart in Japanese?
See the four-tool result →
NEW BUSINESS TESTFive AI tools faced Japanese codes, prices, dates, API, and CSV
See the native-reviewed result →
NEW COMPARISONNotta vs Sonix, tested on the same eight Japanese recordings
Read the full comparison →
NEW ACCURACY GUIDEWhat six real tests revealed about Japanese AI transcription
See the errors to check →
01

MEETING & NOTES

Notta

Current leader

Best contextual Japanese

Strongest overall across eight recordings, but our new workplace homophone test exposed a 制作/製作 meaning error.

4 meaning errors18/18 scripted speaker turns8/8 files completed

Watch out: It removed hyphens from A-2048 and XR-205, wrote numbers in kanji, and rendered 青葉 phonetically as アオバ.

Read the full reviewTry Notta for freeAffiliate link · We may earn a commission at no extra cost to you
02

AUDIO TRANSCRIPTION

TurboScribe

Best formatting

Clean numbers and names

Produced readable numbers and codes, but the noisy two-speaker test exposed serious proper-noun and instruction errors.

4 major error occurrences5/6 files completed8/9 noisy turns

Watch out: In noise, 八潮配送センター became 八代海藻センター and 取扱注意 became 無限使い注意.

03

AUDIO TRANSCRIPTION

Sonix

Code errors

Accurate values, weaker kanji context

Excellent speaker labeling in quiet audio, but controlled noise collapsed every turn into one speaker and changed a quantity.

3 major errors9/9 quiet turnsNo noisy separation

Watch out: In noise, 14箱 became 4箱 inside a correction and every voice was labeled Speaker 1, despite 97.69% very-confident words.

Read the full reviewGet 100 free Sonix minutesAffiliate link · We may earn a commission at no extra cost to you
04

MEETING & NOTES

Otter.ai

Price error

Good names, one serious number mistake

Recognized ミライテック and two of three homophones, but changed a price. Later tests were blocked by the Basic plan's import cap.

1 major error2/5 files completed0 free imports left

Watch out: 1万2,800円 became 12,008百円—a materially different amount if read literally.

Try Otter for freeDirect link · No commission

WEB TRANSCRIPTION

JotMe

Not ranked

No transcript produced

The no-signup web converter accepted all five recordings but failed during processing every time.

0 outputs5 source files triedAccuracy not scored

Watch out: This is a web-tool reliability finding, not an accuracy verdict. The signed-in app may behave differently.

Visit JotMeDirect link · No commission

Biggest lesson so far

Polished text can still be confidently wrong.

Sonix reported 95.83% confidence while changing A-2048 to A 248 and XR-205 to エックス 2005. TurboScribe’s clean-looking output changed 仕様 to 使用, while JotMe failed to return text for five different files. In controlled noise, Sonix also lost speaker separation and TurboScribe changed two important terms. Native review still matters.

Our method

Useful evidence, not a feature-list rewrite.

Every tool receives the same original audio. We separate recognition mistakes from punctuation and formatting, and we do not score a tool that fails to produce a transcript.

Read our full scoring standard →
01Complete

Clean everyday Japanese

A short, naturally paced recording used to check basic content accuracy and punctuation.

02Complete

Numbers, names & homophones

Dates, prices, an address, an invented company name, ChatGPT, and 橋・端・箸 in context.

03Complete

Casual speech & corrections

Fillers, four self-corrections, shortened phrases, times, responsibility, and agenda details.

04In progress

Two speakers

Across quiet and noisy takes, Notta labeled all 18 scripted turns correctly. Noise collapsed Sonix to one speaker; TurboScribe labeled 8/9 noisy turns correctly. TurboScribe's quiet take remains pending.

05Complete

Business Japanese & codes

Honorifics, A-2048, XR-205, quantities, prices, a delivery-date change, API, CSV, and a reply deadline.

Read the full test →
06Complete

Contextual Japanese homophones

Two focused recordings tested everyday actions and workplace language, including 採る, 贈る, 直す, 治す, 制作, and 製作.

WHY THIS LAB EXISTS

Japanese accuracy deserves a Japanese review.

A transcript can look polished to a non-Japanese speaker while getting names, kanji, politeness, or meaning wrong. Native Japanese AI Lab is built in Japan to make those mistakes visible—and help international users choose with confidence.

Some product buttons are affiliate links, and each one is labeled beside the button. We may earn a commission after a qualifying purchase, at no extra cost to the reader. Compensation never changes our test files, scoring method, ranking, or the weaknesses we publish.