Japanese expert data
Train your model on Japan's juiciest data
Expert judgment from licensed Japanese practitioners.
Tokyo / Independently held
A model can be fluent in Japanese and still be wrong about Japan.
Every industry here runs on conventions that were never written down in English, and the people who know them are not the people producing training data.
01 / Coverage
01Foundation models
Instruction tuning, RLHF, Japanese benchmarks
02Robotics & manufacturing
Sensor, image and video from Japanese factory floors
03Speech models
Pitch accent, homophones, dialects beyond Tokyo standard
04Voice agents
Keigo. Get the register wrong and it reads as rude
05Healthcare
Records mixing kanji, katakana drug names, English abbreviations
06Legal
English contract data does not transfer to Japanese structure
07Finance
Yūka shōken hōkokusho, earnings calls, domestic broker research
08Insurance
Policy, claims and underwriting follow domestic convention
09Real estate
Leases, building records, tenant correspondence
10Agriculture
Crop, soil and field imagery from Japanese farms
11Construction
Tight urban sites, different equipment and signage
12Your domain
Name the task →
02 / Quality
You should not have to take our word for the quality.
- Agreement (κ)
- Two qualified reviewers score every example independently, and the figure comes with the batch.
- Rework rate
- Counted on the batch you receive, and reported even when it is high.
- Audit log
- Every example traces back to who wrote it and who checked it.
03 / Method
Nothing here came off the internet.
Practitioners write up their own experience as anonymized cases. Nothing proprietary leaves anyone's building, and there is no way to reproduce this with a translation tool or a synthetic pipeline.
04 / Independence
We are in Tokyo, and no model team owns a piece of us.
Nothing you send here ends up shaping a competitor's roadmap. If your side needs the data to stay in Japan, it stays in Japan.
05 / Pilot