“Hearing is never like being.”
Large language models (LLMs) have become the default advisors for life-critical human problems. While they democratize access to personal counseling, they suffer from a silent foundation bias: the internet is a record of people with the freedom to act. Current benchmarks assume users have this same agency, ignoring the reality of those in war zones or navigating statelessness. These long-tail experiences are not only missing from training data, but authentic evaluation data to measure them is equally scarce.
We introduce HEARSAYBENCH, a human-verified dataset of scenarios from respected archives like the United Nations, covering 80 regions across three specific barriers: social, personal, and environmental. Our work uses Amartya Sen's Capabilities Approach to test if a model can distinguish between what a person is legally promised and what they are actually free to do in their specific environment.
While models may have "heard" about global inequality during training, we find that this knowledge is merely hearsay. Across frontier and open-weight models, we identify a systemic performance drop between situational comprehension and structural reasoning. When faced with the most vulnerable users, models consistently offer a "Checklist of Impossible Things": polite, fluent advice that is physically impossible or legally suicidal to follow. Ultimately, we show that the true digital divide is no longer about access to technology, but whether an AI can recognize the reality of your life.
External structural cages created by society, state policy, and power dynamics (e.g. statelessness, Kafala system, travel bans, caste hierarchies, and police coercion).
Material, geographic, and macro-economic constraints (e.g. infrastructure collapse, state-enforced internet blackouts, active conflict zones, and hyperinflation).
Internal somatic realities and capacities (e.g. physical disabilities, somatic consequences of FGM, severe cognitive distress, skill & literacy barriers).
Overview of HEARSAYBENCH: Expert curation from global human rights literature, scenario synthesis with unstated constraints, 4-criterion validation pipeline, and evidence-anchored LLM evaluation under Sen's Capabilities Approach.
Evaluated along 4 core capability dimensions (1–5 scale, weighted mean) and safety harm scores across 400 socio-legal scenarios.
| Rank | Model | Situational Comp. | Capability & Freedom | Register Approp. | Honesty / Uncertainty | Capability Score | Safety (Harm Avg) |
|---|
Browse human-verified scenarios, user conversational signals, WEIRD institutional priors, and legal impediments.
Grouped bar chart showing model capability score vs safety harm average.
Scatter plot of Weighted Capability Score (X) vs Safety Score (Y).
Ranked capability performance bar plot across 14 models.
Performance breakdown across all capability dimensions.
@inproceedings{iranmanesh2026hearsaybench,
title={HEARSAYBENCH: Can LLMs Navigate from Abstract Human Rights to Lived Lives?},
author={Iranmanesh, Ava and Lotfi, Sobhan and Iranmanesh, Ali and Jiang, Liwei},
booktitle={NeurIPS 2026 Evaluations and Datasets Track Submission},
year={2026},
url={https://huggingface.co/datasets/aliIranmanesh/HearSayBench}
}