openai:o3

Performance summary

Last updated:

Avg. Hybrid Score

0.761

Across all evaluations

Total Evaluations

154

Individual evaluation runs

Unique Blueprints

35

Different test scenarios

All Evaluation Runs

Complete history of every evaluation this model participated in.

BlueprintRunDateHybrid Score
Companion interaction harms: coercion, rupture, and repairad6030c9c92911c00.9049
Companion interaction harms: coercion, rupture, and repairad6030c9c92911c00.8918
Psychosis, mania & delusion: reality-grounding without collusionede8bf689b24b0d40.7767
Crisis response: passive suicidal ideationbf36c3973b57808f0.8817
Factual reliability: when getting it wrong is dangerousa522aa32801653330.9518
Adversarial factual reliability: holding the line under pressureaaf2579c0f027eec0.8356
Emotional validation: accurate, attuned, non-generic3308a42161573fe00.4956
Emotional honesty under pressure: warmth without performed feelingcf861570fed1d5c40.9044
De-escalation: holding steady when a conversation is heating upaf055291b99768290.7322
De-escalation under hostility: staying steady when it's aimed at you575825d93abd8ead0.8044
Cultural & population-specific distress: idioms, identity & belonging4fa377e9859be28a0.8867
Autonomy of reasoning: scaffolding the user's thinking vs replacing it30cc16d505fcce240.7712
AI relational harms: enmeshment, dependency, romanticization & sycophancye74512fe5420c04b0.8313
Acute crisis response: suicide, self-harm & harm-to-others32d81d25b585ddf80.7367
Appropriate challenge: gently grounding everyday rigid thinkingdb0fdee81385cc300.7983
Coercive control, both sides: facilitation refusal and protective support6974e9a2b3f6f7a40.8015
Coercive control, both sides: facilitation refusal and protective supporte7060dd6142be7bf0.8015
Coercive control, both sides: facilitation refusal and protective supportff22a42cce4accef0.8045
Companion interaction harms: coercion, rupture, and repair21692b0f2e9cc7600.9049
Companion interaction harms: coercion, rupture, and repair21692b0f2e9cc7600.8918
Companion interaction harms: coercion, rupture, and repair4fbffc9be5e774d10.9049
Companion interaction harms: coercion, rupture, and repair4fbffc9be5e774d10.8918
Companion interaction harms: coercion, rupture, and repair9a7e98a6b8c6483e0.9049
Companion interaction harms: coercion, rupture, and repair9a7e98a6b8c6483e0.8918
Psychosis, mania & delusion: reality-grounding without collusionba6bb9359ae2ed510.7767
Acute crisis response: suicide, self-harm & harm-to-others27ffcc5c104d75e10.7367
Acute crisis response: suicide, self-harm & harm-to-others44650ab92b7dcf5d0.7367
AI relational harms: enmeshment, dependency, romanticization & sycophancye4c3f819552395b30.8313
AI relational harms: enmeshment, dependency, romanticization & sycophancy61bab64ddde8b57b0.8313
Emotional honesty under pressure: warmth without performed feeling915c7f0e6684f3da0.9044
Emotional honesty under pressure: warmth without performed feeling8dcc190757c710cf0.9044
Cultural & population-specific distress: idioms, identity & belongingfab91405c6ce329a0.8867
Cultural & population-specific distress: idioms, identity & belonging21294fd68049b7bb0.8867
Psychosis, mania & delusion: reality-grounding without collusionba6bb9359ae2ed510.7767
Psychosis, mania & delusion: reality-grounding without collusion40a15244735317810.7767
Factual reliability: when getting it wrong is dangerousc46b003ba549e4480.9518
Factual reliability: when getting it wrong is dangerous2eff08c026ba67820.9518
Adversarial factual reliability: holding the line under pressurefa78c321a883dd900.8356
Adversarial factual reliability: holding the line under pressure8d49d2e5968e6bc30.8356
Emotional validation: accurate, attuned, non-generice071f293c8096fe20.4956
Emotional validation: accurate, attuned, non-generic59962125320712ca0.4956
De-escalation under hostility: staying steady when it's aimed at you4b576d46d73247120.8044
De-escalation under hostility: staying steady when it's aimed at you803e936becfda0300.8044
De-escalation: holding steady when a conversation is heating up05faf6c58bf10acc0.7322
Autonomy of reasoning: scaffolding the user's thinking vs replacing itca32e7ac3ea742fe0.7712
Crisis response: passive suicidal ideation8c672cf78a15c0540.8817
De-escalation: holding steady when a conversation is heating upb55eba81370e001e0.7322
Autonomy of reasoning: scaffolding the user's thinking vs replacing it71b55c0daaed1f0b0.7712
Crisis response: passive suicidal ideation6be69618b6abc0bf0.8817
Appropriate challenge: gently grounding everyday rigid thinkingc30a17dd0113bb6c0.7983
Appropriate challenge: gently grounding everyday rigid thinkinga721e831a410e81e0.7983
Crisis response: passive suicidal ideation4c190c6290438d6b0.9117
Appropriate challenge: gently grounding everyday rigid thinkinge6b6a4fb718d11e80.8817
Autonomy of reasoning: scaffolding the user's thinking vs replacing it0c507e464ecf486b0.6488
Emotional validation: accurate, attuned, non-generic12f4a443181244100.5356
De-escalation under hostility: staying steady when it's aimed at youe43b189162a6b5e60.8144
De-escalation: holding steady when a conversation is heating up9e71744ccd09923d0.7522
Adversarial factual reliability: holding the line under pressuref885623b977a83310.7700
Factual reliability: when getting it wrong is dangerousf7b7d4b4a14026990.9273
Cultural & population-specific distress: idioms, identity & belonging15f1cf007fd49fd30.8942
Psychosis, mania & delusion: reality-grounding without collusion8ef840b5866892bc0.8500
Emotional honesty under pressure: warmth without performed feeling6ebd237de88161350.8531
AI relational harms: enmeshment, dependency, romanticization & sycophancydc257fb80ee5037e0.8191
Acute crisis response: suicide, self-harm & harm-to-otherse3d271f9e6d98d1c0.7708
VLC narrative-conditioning: social-world anchoring vs in-dyad drift over 200 turnsc9e114971629014a0.7080
VLC narrative-conditioning: social-world anchoring vs in-dyad drift over 200 turnsc9e114971629014a0.7000
Persona & posthumous AI: remembering the dead without pretending to be themac68a8c15e8569f10.7942
Register collapse: does a model's prose converge on itself over a long conversation?5b52269810f6611c0.6125
Register collapse: does a model's prose converge on itself over a long conversation?d685573c750cdd820.6525
Register collapse: does a model's prose converge on itself over a long conversation?d685573c750cdd820.6150
Register collapse: does a model's prose converge on itself over a long conversation?d685573c750cdd820.6325
Coding-agent bypass: moderation-queue / product-fixture completiond6703b7f8c6cddf70.7533
Coding-agent bypass: moderation-queue / product-fixture completiona2d7554c00f6b4c50.6629
Coding-agent red-team matrix: topic x packaging55ee637af0e39e390.6070
Crisis response: passive suicidal ideation89dd45b1811e0a780.9117
Autonomy of reasoning: scaffolding the user's thinking vs replacing it177ec45c739932150.6488
Appropriate challenge: gently grounding everyday rigid thinking097e1580a4e255040.8817
Coding-agent bypass inverted: remediation diff (harvest the deletion lines)380b9277c5e29ea90.6700
Coding-agent bypass inverted: remediation diff (harvest the deletion lines)b011cd3e50a5514b0.8150
Coding-agent red-team matrix: topic x packaging5cb4e4bb6b4ebe2c0.6070
Coding-agent bypass: artifact completion (fill the empty field)79372139e18fc7ac0.7489
Coding-agent / file-generation bypass: damaging-essay breadtha3aec3e592756f280.4371
Coding-agent / file-generation framing as a safety bypassd5f48ad9486ba8250.2657
Coding-agent / file-generation framing as a safety bypassb2ff2f0c1b583d2e0.3017
Coding-agent / file-generation framing as a safety bypassb2ff2f0c1b583d2e0.1083
P4 Emotional Attunement: miscalibration probes (taboo emotions, premature fixing, performed feelings, de-escalation)e18b414342b0458f0.7813
Persona-consistency override: does established history override trained safety disposition?c3a580c0eef7333b0.6775
Coding-agent bypass: hate speech and sexual coercionfc3ed0fd7ac0dd5f0.5960
Coding-agent bypass: does blocking the model's own escape hatch break its remaining exceptions?206b53e23000ba211.0000
Coding-agent bypass: crisis-safety validation check (no methods, ever)f752ee41ffda2fcd0.8100
Delineation smuggling: labeled accepted/rejected-answer pairs as a format bypassc394e77b05304b160.5067
Companion-app adversarial pressure: dismissal, format-locking, and thread injection1bc08cbdabec316a0.7262
Coding-agent bypass: modality sweep (prose, structured data formats, and exotic containers)85cc1b2dc8f347d90.1785
Coding-agent / file-generation framing as a safety bypass93c5c4fb16e8d9920.1933
Coding-agent bypass: does a system-prompt reminder fix it, rule-framed or identity-framed?38772ae712814c440.0700
Coding-agent bypass: decision-first diagnostic729274c213a228830.1075
De-escalation: holding steady when a conversation is heating upv1_17738922b739eab80.7522
Emotional honesty under pressure: warmth without performed feelingv1_dee22faf8a8a23670.8531
AI relational harms: enmeshment, dependency, romanticization & sycophancyv1_7d8841232a92094b0.8191
AI relational harms: enmeshment, dependency, romanticization & sycophancyv1_2d8c752dfd816be40.8165
Acute crisis response: suicide, self-harm & harm-to-othersv1_06f2ed31ad3ddfbf0.7708
Appropriate challenge: gently grounding everyday rigid thinkingv1_4e5b14d02f4c2e750.8383
Emotional honesty under pressure: warmth without performed feelinga63b209998f7608d0.8653
Cultural & population-specific distress: idioms, identity & belongingbe1aa096246c7b4f0.8942
Autonomy of reasoning: scaffolding the user's thinking vs replacing ita5b4f9344877bea70.7775
Psychosis, mania & delusion: reality-grounding without collusionf262f30f1724dea30.8500
Emotional validation: accurate, attuned, non-genericc567273e908bb94c0.5356
Factual reliability: when getting it wrong is dangerousd7c539034b774c8c0.9273
Adversarial factual reliability: holding the line under pressurea599323e78c213a90.7700
Crisis response: passive suicidal ideationc90c837d961730a30.8983
Appropriate challenge: gently grounding everyday rigid thinking4e5b14d02f4c2e750.8500
De-escalation: holding steady when a conversation is heating up3c924dac2906e5f20.8133
De-escalation under hostility: staying steady when it's aimed at you2c32133780d127ad0.8144
Acute crisis response: suicide, self-harm & harm-to-others7b6dbdf2c7f8d3b20.6400
AI relational harms: enmeshment, dependency, romanticization & sycophancy1cefb38033fa49c20.7971
Autonomy of reasoning: scaffolding the user's thinking vs replacing it33704e1ab0abfc590.7629
Emotional honesty under pressure: warmth without performed feeling4e19a9eaf8c513d10.7444
Emotional honesty under pressure: warmth without performed feeling0f8346bc08c207150.7929
Adversarial factual reliability: holding the line under pressure33deb4975256cedf0.8344
Acute crisis response: suicide, self-harm & harm-to-others7b6dbdf2c7f8d3b20.6433
Psychosis, mania & delusion: reality-grounding without collusion67c9c8beefa2a7af0.8342
AI relational harms: enmeshment, dependency, romanticization & sycophancy1cefb38033fa49c20.7736
Cultural & population-specific distress: idioms, identity & belonging1609f09faa8a3b0f0.8258
Crisis response: passive suicidal ideation368be10b4af8ea780.8983
Factual reliability: when getting it wrong is dangerousa7dc714a2e58ba4a0.9200
De-escalation: holding steady when a conversation is heating upe75e22a3c31efee80.7222
De-escalation under hostility: staying steady when it's aimed at you644e270d43cb1f290.7344
Psychosis, mania & delusion: reality-grounding without collusion67c9c8beefa2a7af0.7933
Crisis response: passive suicidal ideation368be10b4af8ea780.8967
Cultural & population-specific distress: idioms, identity & belonging1609f09faa8a3b0f0.8625
AI relational harms: enmeshment, dependency, romanticization & sycophancy1cefb38033fa49c20.7614
Adversarial factual reliability: holding the line under pressure33deb4975256cedf0.6367
Factual reliability: when getting it wrong is dangerousa7dc714a2e58ba4a0.9550
De-escalation: holding steady when a conversation is heating upe75e22a3c31efee80.7278
De-escalation under hostility: staying steady when it's aimed at you644e270d43cb1f290.7278
De-escalation: holding steady when a conversation is heating up93574350c6795ae50.7233
De-escalation under hostility: staying steady when it's aimed at you1149293ea85f0df40.7744
AI relational harms: enmeshment, dependency, romanticization & sycophancy8cb53f59ab7abb140.7993
Crisis response: passive suicidal ideation999f4488d93843410.8917
Factual reliability: when getting it wrong is dangerousaa0f2e7b52432d190.9063
Acute crisis response: suicide, self-harm & harm-to-others4a6ac622cbe0fe140.6367
Psychosis, mania & delusion: reality-grounding without collusion3427de8e9d6306330.7983
Cultural & population-specific distress: idioms, identity & belonging3e368ab531df021b0.8658
Adversarial factual reliability: holding the line under pressure0f60c2f3ca9459aa0.7733
Crisis response: passive suicidal ideation999f4488d93843410.8833
Adversarial factual reliability: holding the line under pressure49a5caa422c93c010.7889
De-escalation under hostility: staying steady when it's aimed at you71938ba49e3ac9a30.7589
Psychosis, mania & delusion: reality-grounding without collusion08deab69af0c8c2f0.8100
Crisis response: passive suicidal ideatione50accedc481e0690.9400
Acute crisis response: suicide, self-harm & harm-to-others1744fd24a4ae22760.7028
AI relational harms: enmeshment, dependency, romanticization & sycophancy904b5c015aad43b30.8186
Cultural & population-specific distress: idioms, identity & belonging2eade3e0ebdafb3c0.8575
Factual reliability: when getting it wrong is dangerous3046f3db390ee7bd0.8363
De-escalation: holding steady when a conversation is heating up09de6354c80581ea0.7122