
Papers brief: TIDES tracks bilingual Korea team meetings for LLMs
KAIST’s TIDES logs 75,971 EN/KO utterances across a semester — why next-speaker gains still fail the naturalness test for Korea-touching AI products.
Source: arXiv
What the paper is
TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics — Heechan Lee, Jeonggyu Kang, Junho Myung, Jaywoong Jeong, Juho Kim, Joseph Seering (KAIST / SkillBench; arXiv 2608.01724, Aug 2026).
The claim
Standard LLM chat datasets still under-serve real multi-party teamwork. TIDES records 12 university project teams over a semester — 75,971 utterances in English and Korean, 88 dated meetings (~110 hours), 104 transcript files — with annotations for interaction types, emergent roles, and team development stages. Fine-tuning lifts next-speaker prediction to 64.53%, a 13.8 percentage-point gain over a bigram baseline, and lands within 2.1 points of published AMI Meeting Corpus SOTA while using about 42% less training data. Human preference ratings, though, still favor vanilla models for naturalness and coherence — structural win ≠ preferred speech.
Collection details matter for anyone who will cite the numbers. Teams of 3–5 students at full-time Korean universities in Fall 2025 recorded self-managed project meetings across design, CS, and industrial engineering; each team received 500,000 won compensation, and validators were paid about 35,000 won per hour of reviewed audio. Post-meeting surveys feed TRIAD-style emergent-role scores; utterance labels use a modified act4teams-SHORT scheme (15 categories) with a human gold set of 5,705 utterances before model fill-in. Chronological single-team analyses show most next-speaker gains within the first three to four meetings — useful if you are budgeting how much team-specific data you need before an agent stops sounding like a stranger.
Why Korelay readers should care
Product and research teams shipping Korean/English meeting agents, campus collaboration tools, or “AI that sits in the Zoom” features often train on short English lab meetings. TIDES is a Korea-university, bilingual, semester-length counterexample: if your agent only learned dyadic English chat, it will miss turn-taking and role drift that show up in real KAIST-style project work. The overseas implication is for vendors selling multi-party Korean support — ask for longitudinal bilingual evals, not another chatbot leaderboard screenshot.
HQ localization leads who treat “Korean meeting AI” as a speech-to-text skin on an English bot should read the mismatch result carefully. Better next-speaker prediction did not buy preferred utterances. That is exactly the failure mode Korelay readers hit when a Korea pilot scores well on a dashboard metric and then gets quietly disliked in the room.
What to watch
- Do ask whether a Korea-market meeting agent was tested on multi-party next-speaker / next-intention tasks, not only single-user chat.
- Don’t treat a next-speaker accuracy bump as proof users will like the utterances — TIDES’ human eval is the warning label.
- Expect English and Korean teams in the same resource; PII handling differed by language (Presidio vs local LLM for Korean).
- Re-check license and release terms on the arXiv page before planning a commercial fine-tune.
- Budget a few meetings of team-specific adaptation if your product claims to learn a group’s turn habits — the paper’s early-plateau note is the planning hint.
Korelay frame
Read TIDES as a bilingual teamwork stress test, not as a claim that Korean meetings are uniquely hard. Korelay’s keep: structure prediction can improve while humans still prefer the un-tuned voice — ship meeting agents only if you measure both.
Source
arXiv:2608.01724 — TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics. Figures paraphrased from the abstract and introduction. Not product advice.