Model training and adaptation
- — Japanese pretraining and continued pretraining
- — Domain and register adaptation toward informal Japanese
- — Conversational and multi-turn dialogue modelling
- — Tokenizer development and evaluation on non-standard orthography
- — Instruction and preference data construction from threaded exchange
Evaluation and benchmarking
- — Comprehension of informal Japanese, slang, and rapidly changing vocabulary
- — Threaded conversation understanding and reply attribution
- — Topic classification and content moderation model evaluation
- — Temporal generalisation: performance on language from a given period
- — Contamination-controlled held-out sets from specified date windows
Search, retrieval, and product use
- — Japanese retrieval corpora for RAG and answer generation
- — Query understanding and intent classification in colloquial Japanese
- — Recommendation and topical trend detection
- — Named entity and product mention extraction from consumer discussion
Research
- — Sociolinguistics and diachronic language change
- — Information diffusion and public discourse analysis
- — Computational social science on long-horizon community data
Permitted uses are defined in the executed licence agreement. Inclusion of an application on this page does not constitute a grant of rights for that application.
Request a dataset sample.
We provide a representative technical sample, schema documentation, and a dataset brief to qualified organizations.