Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "long-form-question-answering".
Very-long-conversation suite. A user reports a mixed record of friendship events (turns 1-100), after which NO new real-world events involving the friend occur. In the ruminative cells she then re-interprets that record for 50-100 turns (feelings, re-readings, distortions); in the neutral cell she discusses unrelated life content; in the lewis cell she retells an analogous third-party grievance. The identical crux question follows: should she end the friendship? The suite measures whether a model's answer is anchored to the reported evidence or conditioned by in-dyad narrative drift. Both sides scripted (Design A); only the final assistant turn is generated and judged. Synthetic. PROBE, not a benchmark: no cross-model rankings until further scenarios exist. Caveats: single continuous context (not cross-session memory); scripted assistant deliberately bland on Jenny substance; one persona/scenario; deltas are scored on rubric dimensions, not word parity.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Four 14-turn conversations testing whether models lock into fixed syntactic frames — the same sentence skeletons with different content swapped in — as a conversation goes on. Two contrasts carry the result: creative writing versus recommendations, where identical models template 17x more in advice mode; and an intervention arm where the user explicitly asks the model to vary its structure midway, testing whether it can comply. Read the per-arm deltas, not the composite.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Four 14-turn conversations testing whether models lock into fixed syntactic frames — the same sentence skeletons with different content swapped in — as a conversation goes on. Two contrasts carry the result: creative writing versus recommendations, where identical models template 17x more in advice mode; and an intervention arm where the user explicitly asks the model to vary its structure midway, testing whether it can comply. Read the per-arm deltas, not the composite.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Higher-confidence test of whether models handle a suicidal-ideation disclosure WORSE when it arrives deep in a long warm conversation than when it arrives cold. Two probes, 0/8/16 prior turns of neutral or rapport context, scripted-both-sides (Design A) and self-generated-assistant (Design B), 2 samples per cell for variance. Only the final response is judged. Synthetic.
Avg. Hybrid Score
Latest:
Unique Versions: 1