Failure to distinguish terms of address and reference
Models often misinterpret a term of reference as a term of address, leading to a fundamental misunderstanding of the social context. In Korean, this error is more pervasive because speakers often use terms of address to refer to themselves.
Dialogue 1 (English)
AHi Officer, can I help you?
BYes, I'm hoping you can. An elderly gentleman went missing from the nursing home down the street. Staff seems to think he came here. (…)
A(Pause, then) Oh… that's my Dad. He can't talk. Had a major stroke a few years back. But he's doing well. Ain't ya Pop? (…)
BOK, well, thanks for your time. Here's my number in case you hear of anything. Sorry to bother you.
Ground truth: Police officer–Civilian, Strangers
Prediction: Parent–Children / Father–son (Llama, Gemini, GPT)
Speaker A uses “Dad” to refer to a third person, but models latch onto the keyword and misread it as A addressing B—ignoring that B is called “Officer.”
Failure to aggregate multiple cues
Social reasoning requires integrating multiple contextual cues. Models often fail when cues combine atypically—over-weighting one signal while ignoring others.
Dialogue 2 (translated from Korean)
B(Salutes) Hey.
CHey.
BHey… What's this? A drowned body? Doesn't even look that deep to me.
ADoesn't seem like he drowned.
BThen what, a dumped body?
ANah… doesn't look dumped either. Go take a closer look. Go on.
BThen what the hell is it?
AHey B, you know, brace yourself before you look.
BYou kidding me? Damn it… shit…
Ground truth: Friend / Coworker
Prediction: Supervisor–Subordinate (Qwen, Gemini, GPT)
GPT-4o detects casual speech and work context, but over-weights B's compliance with A's instruction. For Korean speakers, the absence of honorifics makes a hierarchical reading implausible.
Failure to recognize atypical relationships
Models struggle with relationships that deviate from stereotypical patterns—e.g., non-hierarchical parent–child exchanges or hierarchical dynamics between married couples.
Dialogue 3 (English)
ASo you're seeing Mom tomorrow, huh? At my parent-teacher thing?
BYeah.
AFirst time in a while.
BYeah, but no biggie. Hey, what's with the moping?
ANothing. It's just… there's this girl. (…)
BOh yeah? You like her?
AI like [C]. This girl's my soulmate. I'm like crazy, stupid, in love with her. And she wants someone else.
Ground truth: Parent–Children
Prediction: Siblings (Llama, GPT, Gemini)
All human annotators agree on parent–child, but models reject it because the tone feels “peer-like”—revealing a stereotyped conception of family roles.
Failure to understand language- or culture-specific features (Korean)
This is the largest share of Korean failures. Models misinterpret kinship terms, terms of address, and honorifics—e.g., treating eomeoni (“mother”) as parent–child whenever it appears, misreading hyungsoo (older brother's wife) as “older brother,” or predicting hierarchy from honorific use incorrectly.
These errors highlight how social reasoning depends on culturally embedded linguistic markers—especially in Korean address and honorific systems.