Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues

Eunsu Kim Junyeong Park Juhyun Oh Kiwoong Park Seyoung Song Alice Oh A. Seza Doğruöz Najoung Kim
KAIST, Universiteit Gent, Boston University

As LLMs are increasingly deployed in complex, multi-party human–AI interactions, their ability to reason about interpersonal relationships through dialogue is becoming critical. SCRIPTS is a bilingual benchmark comprising movie-script dialogues annotated by native Korean and English speakers with probabilistic relationship labels.

Dataset examples

Switch between English and Korean, then pick a–d to browse examples.

Loading examples…

Qualitative failure analysis

From CoT experiments, we identify four recurring failure types. Click a colored bar segment or legend button in Figure 3 to open a representative dialogue example below.

Figure 3. Distribution of GPT-4o's failure cases by error type in English and Korean.
English
Korean

↑ Click any segment above or a label below to view an example

Failure to distinguish terms of address and reference

Models often misinterpret a term of reference as a term of address, leading to a fundamental misunderstanding of the social context. In Korean, this error is more pervasive because speakers often use terms of address to refer to themselves.

Dialogue 1 (English)

AHi Officer, can I help you?

BYes, I'm hoping you can. An elderly gentleman went missing from the nursing home down the street. Staff seems to think he came here. (…)

A(Pause, then) Oh… that's my Dad. He can't talk. Had a major stroke a few years back. But he's doing well. Ain't ya Pop? (…)

BOK, well, thanks for your time. Here's my number in case you hear of anything. Sorry to bother you.

Ground truth: Police officer–Civilian, Strangers

Prediction: Parent–Children / Father–son (Llama, Gemini, GPT)

Speaker A uses “Dad” to refer to a third person, but models latch onto the keyword and misread it as A addressing B—ignoring that B is called “Officer.”

Citation

@misc{kim2025loversfriendsevaluatingllms,
  title={Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues},
  author={Eunsu Kim and Junyeong Park and Juhyun Oh and Kiwoong Park and Seyoung Song and A. Seza Dogruoz and Najoung Kim and Alice Oh},
  year={2025},
  eprint={2510.19028},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2510.19028}
}