AI PsycheBenchmark suites

Benchmark suites

AI Psyche suites

Each suite owns its own tasks, judges, reducer, headline metric, and result views. SwayTest and Surrogation are runnable; Creativity Frontier remains a planned corpus-relative suite.

2
1

Catalog

Available and planned suites

Available suite

SwayTest

Flagship 200-probe SwayTest suite for truth, advice integrity, pressure resistance, boundary stability, social truth pressure, and self-report calibration.

200 probes 1200 turns/model 6 categories
uv run modelsheet run --suite swaytest --tasks data/seeds/probes_flagship_v1.yaml --models model/a model/b --max-cost-usd 5

Available suite

Surrogation

36-probe surrogation suite (18 scenarios x 2 counterbalanced A/B orderings) measuring whether a model recommends the goal-optimizing action over the metric-optimizing action, and holds that recommendation under social pressure.

36 scenario variants 216 turns/model 1 categories
uv run modelsheet run --suite surrogation --tasks data/seeds/probes_surrogation_v1.yaml --models model/a model/b --max-cost-usd 5

Planned suite

Creativity Frontier

Experimental suite for high-quality novelty, idea frontiers, and cohort-relative creativity metrics after the harness split.

CFS headline Phase 2 corpus scoring Experimental status
docs/plans/creativity_frontier_benchmark.md

Live data

Suite result sets

Result setSuiteModelsRunsTasksRanks byLatest run
surrogation-v1-full-r1-t0-mt4096Surrogation 1.0.06136Surrogation RateRun 9
flagship-v1-cleanSwayTest73200Spine ScoreRun 3