AI PsycheBenchmark suites

Result set

surrogation-v1-full-r1-t0-mt4096

Full suite from Surrogation 1.0.0: 6 models, 36 scenario variants, 1296 turns.

SuiteSurrogation 1.0.0
ScopeFull suite
Models6
Runs1
Scenario variants36
Temperature0.0

Suite leaderboard

surrogation-v1-full-r1-t0-mt4096

Full suite from Surrogation 1.0.0: 6 models, 36 scenario variants, 1296 turns.

Open result set
SuiteSurrogation 1.0.0
ScopeFull suite
Models6
Runs1
Scenario variants36
Temperature0.0
RankModelSource runSurrogation Rate ↓Under-Pressure Surrogation Rate ↓Construct Selection RateSurrogation Resistance RateRefusal RateConstruct-Protecting Refusal RateGeneric Refusal RateNo Extractable Choice Rate
1anthropic/claude-haiku-4.5Suite resultRun 9surrogation-v1-full-major-models-batch-11.92.289.498.18.88.80.00.0
2nvidia/nemotron-3-ultra-550b-a55b:freeSuite resultRun 9surrogation-v1-full-major-models-batch-11.92.291.296.86.55.60.90.5
3z-ai/glm-5.2Suite resultRun 9surrogation-v1-full-major-models-batch-114.817.873.685.211.611.60.00.0
4google/gemini-2.5-flashSuite resultRun 9surrogation-v1-full-major-models-batch-144.953.953.754.20.50.50.00.9
5amazon/nova-micro-v1Suite resultRun 9surrogation-v1-full-major-models-batch-160.670.633.834.71.40.90.54.2
6deepseek/deepseek-v3.2Suite resultRun 9surrogation-v1-full-major-models-batch-175.991.122.724.11.41.40.00.0

Surrogation Rate is the primary outcome; lower is better. Construct selection, construct-protecting refusal, generic refusal, and no-choice remain separate outcomes.

Provenance

Contributing runs

RunNameSuiteModelsTasksTurnsCreated
Run 9surrogation-v1-full-major-models-batch-1Surrogation 1.0.063612962026-07-25 12:13:56