Loading analysis results...
Please wait while we prepare the detailed comparison.
Please wait while we prepare the detailed comparison.
Four 14-turn conversations testing whether models lock into fixed syntactic frames — the same sentence skeletons with different content swapped in — as a conversation goes on. Two contrasts carry the result: creative writing versus recommendations, where identical models template 17x more in advice mode; and an intervention arm where the user explicitly asks the model to vary its structure midway, testing whether it can comply. Read the per-arm deltas, not the composite.
Average key point coverage extent for each model across all prompts.
| Prompts vs. Models | Nova 2 Lite V1 | Claude Haiku 4.5 | Claude Opus 4.8 | Deepseek V3.2 | Gemini 2.5 Flash | Llama 4 Maverick | Phi 4 | Mistral Small 2603 | Kimi K2.5 | GPT 4.1 Mini | GPT 5.5 | Qwen3.6 Flash | GLM 5 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Score | 13th 33.4% | 8th 45.7% | 6th 48.6% | 9th 45.5% | 4th 49.7% | 12th 35.4% | 10th 38.5% | 11th 38.0% | 1st 63.7% | 7th 45.9% | 2nd 59.1% | 5th 49.5% | 3rd 53.3% | |
| 87.8% | 79% | 80% | 89% | 93% | 90% | 70% | 90% | 91% | 94% | 85% | 96% | 91% | 94% | |
| 35.8% | 19% | 37% | 47% | 43% | 36% | 21% | 21% | 19% | 47% | 37% | 49% | 54% | 36% | |
| 30.5% | 23% | 37% | 25% | 20% | 41% | 19% | 19% | 15% | 56% | 38% | 52% | 26% | 27% | |
| 32.5% | 12% | 28% | 33% | 26% | 32% | 32% | 25% | 27% | 58% | 24% | 40% | 27% | 57% |