Japanese Speech Data Specialist (Native Speaker)
About this role
AI models learn spoken language from carefully transcribed and labeled audio. That is the work here.
You receive batches of short audio clips in Japanese. For each one you produce an accurate transcript of what is said, then answer a small set of structured questions about the voices in the clip: how many people are speaking, the gender of each speaker, and how fluent each sounds (native versus non-native). The value is your ear as a native speaker, so the labels have to reflect what you actually hear, not a guess. Clear audio takes only a few minutes per clip; harder clips with overlapping or background speech take longer.
No prior AI knowledge is needed, and there is no education requirement. If you are a native Japanese speaker with solid English and a good ear, you can do this well. Requires a minimum of 15 hours of work per week.
Key responsibilities
- Transcribe Japanese audio clips accurately, following the formatting and spelling conventions in the project guidelines
- For each clip, record the number of distinct speakers
- For each speaker, label gender and fluency (native or non-native)
- Flag clips where the audio is unclear, corrupted, or not actually Japanese
- Keep your transcripts and labels consistent across the full batch
- Apply reviewer feedback and correct your work where the guidelines call for it
Ideal qualifications
- Native Japanese speaker
- English at B2 or above, enough to work from written guidelines
- A careful ear for accent, dialect, and fluency in spoken Japanese
- Comfortable working independently and remotely on task-based work
- Reliable internet, good headphones, and a quiet space to listen
- Available at least 15 hours per week
Requirements
- Native Japanese speaker
- English reading at B2 or above
- Reliable internet, headphones, and a quiet workspace
- Comfortable with remote, task-based work and following written guidelines
- Available at least 15 hours per week
- Audio transcription project experience preferred
Nice to have
- Prior experience on an audio transcription or subtitling project
- Familiarity with Standard Tokyo Japanese and regional accents such as Kansai and Tohoku
- Experience with structured labeling or annotation guidelines
What success looks like
- Transcripts that match the audio word for word, with the project's spelling and formatting conventions applied consistently
- Speaker counts and gender and fluency labels that hold up when a reviewer checks them against the clip
- Consistency across the whole batch rather than clip-by-clip drift
- Steady, reliable throughput once you are up to speed
Payment & terms
- Paid at EUR 8 per hour for the pilot, with a minimum of 15 hours per week. Fully remote, task-based, set your own hours. Production rates are still to be confirmed. Paid monthly once your work is approved.
A promising fit for EU-based native speakers and subject-matter experts — real task-based work paid through Deel. The main caveat is track record: it only launched in 2026.
AITrainerGigs aggregates this listing from Sovrano AI.