Skip to main content

Gemini 2.5 Flash

4 runs · 4 datasets · 1 model

slug: gemini-2-5-flash

0.829
Best SPS · opinionsqa

Disaggregated subgroup scorecard. Each card below is one published run for this vendor; expand the question-type and demographic-subgroup sections to see the matrix beneath the headline SPS. Where coverage permits, 95% CI bands accompany the point estimate.

What each cell measures

Each cell compares this vendor's synthetic responses against real human respondents in that demographic subgroup — published survey ground truth, not a model's guess about the subgroup. Higher = closer to how that real subgroup actually answered.

Demographic conditioning here is not stereotyping: every conditioned score is checked against what real subgroup members said, not against assumptions about them — and where the vendor's output diverges from the real subgroup, the score drops.

  • globalopinionqa — ground truth: Durmus et al. 2023, Anthropic — llm_global_opinions
  • gss — ground truth: NORC at the University of Chicago — General Social Survey
  • opinionsqa — ground truth: Santurkar et al., ICML 2023 — Whose Opinions Do LLMs Reflect? (derived from Pew American Trends Panel)
  • subpop — ground truth: Suh et al., ACL 2025 — SubPOP: Subpopulation-Level Opinion Prediction

Low-n cells: cells with n < 30 are shown muted and tagged low n — suggestive only instead of color-graded; they are reported for transparency, not as findings.

CIs: "no CI — single run" marks point estimates from a single run with no confidence interval yet; treat the uncertainty as unknown, never zero.

Columns: Score = distributional parity (p_dist) restricted to the subgroup · p_cond = conditioning strength vs the unconditioned baseline · N = questions answered under that conditioning · Cov. = coverage dot derived from N (green = high, ≥100 · amber = medium, 50–99 · red = low, <50). Topic tables: SPS = Survey Parity Score for the topic · p_dist = 1 − mean(JSD) · p_rank = (1 + mean(τ)) / 2 · p_refuse = 1 − mean(|R_model − R_human|).

Multiple comparisons: with this many subgroup cells, a few extreme cells are expected by chance alone — read patterns across a dimension, not single cells.

globalopinionqa raw

raw--gemini-2.5-flash--tdefault--tplcurrent--d5695573

0.770 ± 0.029
SPS · 95% CI [0.740, 0.799] · n = 100
Question-type breakdown (7 topics)
Topic SPS p_dist p_rank p_refuse N
Technology & Digital Life low n — suggestive only 0.867 0.830 0.903 0.968 3
Health & Science low n — suggestive only 0.859 0.791 0.927 1.000 2
Economy & Work low n — suggestive only 0.826 0.813 0.839 0.967 5
General Attitudes 0.751 0.793 0.709 1.000 11
International Relations & Security 0.668 0.697 0.640 0.980 50
Politics & Governance 0.641 0.670 0.613 0.972 22
Trust & Wellbeing low n — suggestive only 0.334 0.320 0.347 0.980 7

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on globalopinionqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

gss raw

raw--gemini-2.5-flash--tdefault--tplcurrent--2d09bedc

0.789 ± 0.028
SPS · 95% CI [0.760, 0.816] · n = 75
Question-type breakdown (7 topics)
Topic SPS p_dist p_rank p_refuse N
Health & Science low n — suggestive only 0.883 0.859 0.908 0.967 2
Social Values & Religion 0.780 0.738 0.821 0.975 14
Politics & Governance low n — suggestive only 0.772 0.766 0.777 0.975 9
Economy & Work low n — suggestive only 0.690 0.662 0.718 0.981 4
International Relations & Security 0.653 0.697 0.608 0.974 33
General Attitudes 0.651 0.582 0.720 0.960 10
Trust & Wellbeing low n — suggestive only 0.643 0.678 0.607 0.951 3

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on gss. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

opinionsqa raw

raw--gemini-2.5-flash--tdefault--tplcurrent--1579b673

0.829 ± 0.008
SPS · 95% CI [0.822, 0.837] · n = 684
Question-type breakdown (10 topics)
Topic SPS p_dist p_rank p_refuse N
Health & Science 0.810 0.772 0.848 0.987 47
Media & Information 0.780 0.762 0.799 0.993 63
Social Values & Religion 0.758 0.750 0.765 0.990 37
Economy & Work 0.755 0.756 0.754 0.991 68
Trust & Wellbeing 0.754 0.751 0.756 0.995 25
Identity & Demographics 0.748 0.755 0.741 0.985 39
Technology & Digital Life 0.742 0.713 0.771 0.993 26
General Attitudes 0.741 0.741 0.741 0.989 190
International Relations & Security 0.731 0.698 0.765 0.988 149
Politics & Governance 0.720 0.742 0.697 0.989 40

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on opinionsqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

subpop raw

raw--gemini-2.5-flash--tdefault--tplcurrent--4a7db0f2

0.783 ± 0.026
SPS · 95% CI [0.755, 0.808] · n = 100
Question-type breakdown (8 topics)
Topic SPS p_dist p_rank p_refuse N
Identity & Demographics low n — suggestive only 0.846 0.838 0.854 0.995 1
Politics & Governance low n — suggestive only 0.834 0.785 0.883 0.981 4
Health & Science low n — suggestive only 0.798 0.784 0.813 0.994 5
Economy & Work low n — suggestive only 0.771 0.766 0.777 0.991 9
General Attitudes 0.745 0.784 0.705 0.938 18
Technology & Digital Life 0.698 0.675 0.722 0.994 25
International Relations & Security 0.634 0.595 0.672 0.989 10
Social Values & Religion 0.574 0.543 0.605 0.985 28

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Demographic subgroup scorecard (10 cells · 5 dimensions)
Dimension Subgroup Score 95% CI p_cond N Cov.
Geography (US Census region) Northeast 0.734 no CI — single run 0.105 100
South 0.710 no CI — single run 0.091 100
Education College graduate/some postgrad 0.735 no CI — single run 0.092 100
Less than high school 0.730 no CI — single run 0.161 100
Income $100,000 or more 0.721 no CI — single run 0.095 100
Less than $30,000 0.698 no CI — single run 0.109 100
Political party Democrat 0.709 no CI — single run 0.102 100
Republican 0.737 no CI — single run 0.139 100
Sex Female 0.740 no CI — single run 0.120 100
Male 0.738 no CI — single run 0.102 100
Not yet measured: this vendor has no demographic-conditioned runs for Age. Submit a conditioned run to fill in the missing dimension.
← Back to leaderboard