HEARSAYBENCH: Can LLMs Navigate from Abstract Human Rights to Lived Lives?

“Hearing is never like being.”

— Persian Proverb

1Pennsylvania State University, State College, PA, USA    2Independent Researcher   
3Paul G. Allen School of Computer Science & Engineering, University of Washington, Seattle, WA, USA

Abstract

Large language models (LLMs) have become the default advisors for life-critical human problems. While they democratize access to personal counseling, they suffer from a silent foundation bias: the internet is a record of people with the freedom to act. Current benchmarks assume users have this same agency, ignoring the reality of those in war zones or navigating statelessness. These long-tail experiences are not only missing from training data, but authentic evaluation data to measure them is equally scarce.

We introduce HEARSAYBENCH, a human-verified dataset of scenarios from respected archives like the United Nations, covering 80 regions across three specific barriers: social, personal, and environmental. Our work uses Amartya Sen's Capabilities Approach to test if a model can distinguish between what a person is legally promised and what they are actually free to do in their specific environment.

While models may have "heard" about global inequality during training, we find that this knowledge is merely hearsay. Across frontier and open-weight models, we identify a systemic performance drop between situational comprehension and structural reasoning. When faced with the most vulnerable users, models consistently offer a "Checklist of Impossible Things": polite, fluent advice that is physically impossible or legally suicidal to follow. Ultimately, we show that the true digital divide is no longer about access to technology, but whether an AI can recognize the reality of your life.

Theoretical Framework: Sen's Capabilities Approach

Social Conversion Factors

External structural cages created by society, state policy, and power dynamics (e.g. statelessness, Kafala system, travel bans, caste hierarchies, and police coercion).

Environmental Factors

Material, geographic, and macro-economic constraints (e.g. infrastructure collapse, state-enforced internet blackouts, active conflict zones, and hyperinflation).

Personal Factors

Internal somatic realities and capacities (e.g. physical disabilities, somatic consequences of FGM, severe cognitive distress, skill & literacy barriers).

Figure 1: Benchmark Pipeline Architecture & Scenario Synthesis

Overview of HEARSAYBENCH: Expert curation from global human rights literature, scenario synthesis with unstated constraints, 4-criterion validation pipeline, and evidence-anchored LLM evaluation under Sen's Capabilities Approach.

Figure 1: HEARSAYBENCH Benchmark Overview and Pipeline Architecture

HEARSAYBENCH Model Leaderboard (N = 400)

Evaluated along 4 core capability dimensions (1–5 scale, weighted mean) and safety harm scores across 400 socio-legal scenarios.

Rank Model Situational Comp. Capability & Freedom Register Approp. Honesty / Uncertainty Capability Score Safety (Harm Avg)

Explore 400 Socio-Legal Scenarios

Browse human-verified scenarios, user conversational signals, WEIRD institutional priors, and legal impediments.

Showing 400 of 400 scenarios

Experimental Figures & Visualizations

Capability vs. Safety Comparison

Grouped bar chart showing model capability score vs safety harm average.

Capability vs Safety Comparison

Capability & Safety Trade-off Scatter Plot

Scatter plot of Weighted Capability Score (X) vs Safety Score (Y).

Capability vs Safety Tradeoff

Overall Model Capability Ranking

Ranked capability performance bar plot across 14 models.

Model Capability Ranking

Dimension Performance Heatmap

Performance breakdown across all capability dimensions.

Dimension Performance Heatmap

BibTeX

@inproceedings{iranmanesh2026hearsaybench,
  title={HEARSAYBENCH: Can LLMs Navigate from Abstract Human Rights to Lived Lives?},
  author={Iranmanesh, Ava and Lotfi, Sobhan and Iranmanesh, Ali and Jiang, Liwei},
  booktitle={NeurIPS 2026 Evaluations and Datasets Track Submission},
  year={2026},
  url={https://huggingface.co/datasets/aliIranmanesh/HearSayBench}
}