Academic papers

Papers brief: TIDES tracks bilingual Korea team meetings for LLMs

KAIST’s TIDES logs 75,971 EN/KO utterances across a semester — why next-speaker gains still fail the naturalness test for Korea-touching AI products.

  • papers
  • llm
  • korean
  • meetings

Source: arXiv

What the paper is

TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics — Heechan Lee, Jeonggyu Kang, Junho Myung, Jaywoong Jeong, Juho Kim, Joseph Seering (KAIST / SkillBench; arXiv 2608.01724, Aug 2026).

The claim

Standard LLM chat datasets still under-serve real multi-party teamwork. TIDES records 12 university project teams over a semester — 75,971 utterances in English and Korean, 88 dated meetings (~110 hours), 104 transcript files — with annotations for interaction types, emergent roles, and team development stages. Fine-tuning lifts next-speaker prediction to 64.53%, a 13.8 percentage-point gain over a bigram baseline, and lands within 2.1 points of published AMI Meeting Corpus SOTA while using about 42% less training data. Human preference ratings, though, still favor vanilla models for naturalness and coherence — structural win ≠ preferred speech.

Collection details matter for anyone who will cite the numbers. Teams of 3–5 students at full-time Korean universities in Fall 2025 recorded self-managed project meetings across design, CS, and industrial engineering; each team received 500,000 won compensation, and validators were paid about 35,000 won per hour of reviewed audio. Post-meeting surveys feed TRIAD-style emergent-role scores; utterance labels use a modified act4teams-SHORT scheme (15 categories) with a human gold set of 5,705 utterances before model fill-in. Chronological single-team analyses show most next-speaker gains within the first three to four meetings — useful if you are budgeting how much team-specific data you need before an agent stops sounding like a stranger.

Why Korelay readers should care

Product and research teams shipping Korean/English meeting agents, campus collaboration tools, or “AI that sits in the Zoom” features often train on short English lab meetings. TIDES is a Korea-university, bilingual, semester-length counterexample: if your agent only learned dyadic English chat, it will miss turn-taking and role drift that show up in real KAIST-style project work. The overseas implication is for vendors selling multi-party Korean support — ask for longitudinal bilingual evals, not another chatbot leaderboard screenshot.

HQ localization leads who treat “Korean meeting AI” as a speech-to-text skin on an English bot should read the mismatch result carefully. Better next-speaker prediction did not buy preferred utterances. That is exactly the failure mode Korelay readers hit when a Korea pilot scores well on a dashboard metric and then gets quietly disliked in the room.

What to watch

  • Do ask whether a Korea-market meeting agent was tested on multi-party next-speaker / next-intention tasks, not only single-user chat.
  • Don’t treat a next-speaker accuracy bump as proof users will like the utterances — TIDES’ human eval is the warning label.
  • Expect English and Korean teams in the same resource; PII handling differed by language (Presidio vs local LLM for Korean).
  • Re-check license and release terms on the arXiv page before planning a commercial fine-tune.
  • Budget a few meetings of team-specific adaptation if your product claims to learn a group’s turn habits — the paper’s early-plateau note is the planning hint.

Korelay frame

Read TIDES as a bilingual teamwork stress test, not as a claim that Korean meetings are uniquely hard. Korelay’s keep: structure prediction can improve while humans still prefer the un-tuned voice — ship meeting agents only if you measure both.

Source

arXiv:2608.01724 — TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics. Figures paraphrased from the abstract and introduction. Not product advice.