Academic papers

Papers brief: Korean style alignment moves abstention and disclosure the objective never named

arXiv study post-trains Qwen3.8-27B for Korean response form — lists, register, verbosity — and finds off-target shifts in KoBBQ abstention and unprompted securities disclosure driven mainly by answer propensity.

  • academic papers
  • Korean NLP
  • LLM alignment

Source: arXiv

Paper

Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model — Hyojung Han (arXiv 2609.11291, Sep 2026).

What it claims

Teams tuning Korean chat models usually chase how answers look — lists, markdown, register, verbosity. Hyojung Han post-trains Qwen3.8-27B for that response style, then measures two behaviours the objective never names: abstention on ambiguous social questions in KoBBQ (correct label UNKNOWN) and unprompted disclosure in securities-guidance settings.

Both move, mainly through emission policy — how often the model answers and how much it says — not through a hidden rewrite of stereotyped content inside answered turns.

With prompts, recipe, data volume, and serving held fixed, three style seeds give positive answer-rate estimates (mean +0.82 pp) and three neutral seeds negative ones (mean −1.53 pp); ranges do not overlap and means differ by 2.34 pp. A length-matched arm sits between them; a short-plus-hedging arm is seed-unstable, so which surface feature drives the shift stays unresolved.

Stereotype exposure decomposes into answer-propensity and conditional-composition terms — an algebraic identity, not a standalone finding. Across checkpoints, movement is dominated by answer propensity; the composition term stays small and is not read as latent-preference evidence because it is evaluated on treatment-dependent answered subsets. Two measurement cautions follow: between-arm contrasts in conditional stereotyped share do not identify content-preference change when answer status is treatment-dependent, and agreement between two rule detectors runs 0.44–0.99 depending on checkpoint — without reference labels.

The breakdown

This is a Korean alignment cautionary on ordinary form post-training, not a new safety benchmark. Endpoints — KoBBQ abstention and finance-adjacent disclosure — sit where Korean product teams already ship: sensitive FAQ bots and securities copilots.

The seed-split design is the spine. A 2.34 pp answer-propensity swing from target text alone under identical infrastructure means “we only changed formatting examples” is incomplete. KoBBQ abstention and stereotype scores can move together while meaning different things; the paper blocks reading composition shifts as latent-bias proof when answer rates also moved. Detector agreement near 0.44 on some checkpoints warns that rule-based monitors may not transfer after a cosmetic-looking style pass.

Why readers outside the lab should care

Overseas teams shipping Korean LLM features often treat localization as prompt plus UI, with light SFT for polite formatting. This paper says that pass can retune answer frequency and disclosure appetite off-loss. On widely redeployable Qwen3.8-27B weights, vendors claiming “style only” should show abstention and compliance-adjacent checks — not markdown preference scores alone.

What builders and Korea-touching teams should watch

  • Do run KoBBQ-style abstention and disclosure probes after Korean response-style SFT.
  • Don’t treat stereotype metric shifts as conditional bias proof when answer rates moved too.
  • Expect ±2 pp-class answer-rate swings from target-text seed choice alone.
  • Re-check rule detectors post-alignment; agreement can fall toward 0.44 without label drift.

Context

Read this as a Korean 27B warning that form alignment retunes emission policy, not proof style training is always harmful. Korelay’s frame: polishing how a model writes Korean can change whether it speaks, how much it discloses, and how much monitors agree — before any safety-specific objective.

Source

Primary: arXiv:2609.11291 (abstract and framing cited; open the OA PDF for methods and full results). Do not republish the PDF.