Results · grounded in the experiment harness
What is settled, and where the frontier is.
Four pieces of the program are done: the framing, a reproducible definition, a metric ablation, and a calibrated negative result. The in-generation controller is the open frontier, where the first coherence-preserving rhythm dial just cleared the noise floor.
Every number on this page comes from the experiment harness. Effect sizes from tiny base models (distilgpt2, gpt2-medium) are directional, not population claims, and we say so where it matters. The transferable results are the ranking of statistics, the noise calibration, and the existence of a monotone dial, not the absolute magnitudes.
A reproducible definition
The burstiness vector.
The field's operational definition of burstiness is a detector blog post: a single scalar contrasting perplexity variation. We replace it with a decomposable, reproducible target that a controller can steer one dimension at a time.
Variance and kurtosis of sentence length, mean surprisal and local surprisal jumpiness under a fixed reference model, and the Shannon entropy of the punctuation pattern. Each component degenerates toward a small value for rhythmically flat text. On the reference demo, a human-like sample scores var(L) about 260 against about 0.49 for a flat sample, and punctuation entropy 2.00 against 0.0.
Metric ablation
The field's implicit metric is the weakest one.
Detection practice operationalizes burstiness as the standard deviation of surprisal. We ablated five surprisal statistics on two sets: a hard 4-versus-4 hand-written dev set, and a 12-versus-12 corpus of public-domain literary prose against distilgpt2 output. The ranking is the result, and it is stable across both sets.
| Statistic | Cohen's d (dev) | Cohen's d (corpus) | Verdict |
|---|---|---|---|
| mean_surprisal | +1.43 | +2.84 | Strongest separator |
| fluc_abs_diff (local jumpiness) | +1.25 | +3.15 | Best fluctuation variant |
| fluc_raw (stdev of surprisal) | +0.59 | +2.19 | Weak on the hard set |
| fluc_windowed | -0.84 | +1.16 | Reverses on dev |
| fluc_cv | -0.52 | -0.90 | Reverses on both, drop it |
Standard deviation of surprisal, the metric the field runs on, is the weakest discriminator on the hard dev set (d = 0.59, medium). Local jumpiness (mean absolute consecutive surprisal difference) and mean surprisal separate human from machine far better. The coefficient of variation reverses sign on both sets, so flat text scores higher: we drop it. The corpus effect sizes are inflated by the easy human-versus-distilgpt2 contrast; the transferable finding is the ordering, not the magnitudes. This is what fixes the vector to B = [var(L), kurt(L), mean_surprisal, fluc_abs_diff, punct_entropy].
A calibrated negative result
var(L) is noise-dominated on short generations.
The first steering arm appeared to work: a single content-matched activation vector seemed to raise var(L) from 84.7 to 137.8, a 63 percent gain at constant coherence. It did not replicate. Across five sampling seeds at the same configuration, var(L) ranged from 4.5 to 61.7, all at or below baseline. The apparent win was a sampling artifact.
The sampling variance of an estimated variance falls only as one over the sentence count. On 4 prompts times 50 tokens you have a handful of sentences, so the standard deviation of var(L) equals or exceeds its mean at every steering scale. Scaling the base model 4x to gpt2-medium did not fix it: var(L) stayed non-monotone with its std at or above its mean. The bottleneck is the estimator, not model capacity.
Single-run steering claims for var(L) are false positives. Any burstiness controller result must be reported with multi-seed averaging and a reported var(L) standard deviation, and evaluation length must clear a minimum sentence count so the estimator settles. Surfacing this before investing in training is exactly what the design gate exists to do. It also rescues every later arm, because the paired long-form protocol it forces is what finally made a real effect visible.
The in-generation controller
A boundary-steered sentence-length dial.
Token-level arms (a single activation vector, a best-of-N LoRA, GRPO) all hit the same wall: they bought variance by pushing the model off-distribution, degrading coherence, and never moved var(L) above its own noise. The fix was altitude. Steer only the sentence-ending punctuation logits toward a synthesized length plan, training-free, with the computable metric as a running discriminator. This controls var(L) at its native altitude, the boundary decision, instead of the token.
Under the paired long-form protocol the negative result forced, the dial moves. The first run, at lambda 8 over 300-token generations across 12 common-random-number pairs, gives a paired difference of +56.4 plus or minus 26.9 in realized var(L) between the low and high plans, clearing twice its standard error, while coherence is unchanged (mean surprisal 2.761 versus 2.770). Every prior arm had raised mean surprisal to buy variance; this one does not.
| Criterion | Result | Status |
|---|---|---|
| C1 monotonicity (Spearman rho over 4 dial levels) | rho = 1.00 | Pass |
| C2 effect, robust direction (sign / Wilcoxon) | 40/48 pairs, p ≤ 3×10-6; dz = 0.64 (magnitude) | Pass |
| C3 coherence drift (mean surprisal) | 0.13 | Pass |
| C4 content (cross-dial vs within-level floor) | 0.36 ≥ 0.29 floor, p = 0.002 | Pass |
At spec-grade scale on gpt2-medium (lambda 8, 480-token generations, 30 sentences, 8 seeds, 4 dial levels, 48 matched pairs), expected var(L) rises monotonically with the dial, Spearman rho = 1.00, and coherence holds (drift about 0.13). The verdict statistic matters: the dial is a within-sample manipulation (low and high share the seed and prompt), so the right test is paired, and var(L) is heavy-tailed, so a single outlier generation swamps any mean-based effect size. That is exactly why Cohen's d wobbled between 0.59 and 0.87 across re-runs: it is the wrong statistic for this metric, an instance of the project's own noise theorem applied to its own verdict. Under the correct rank-based tests the dial is robustly controllable: 40 of 48 matched pairs respond in the commanded direction (sign test p = 3×10-6, Wilcoxon signed-rank z = 4.72, p = 2×10-6, mean effect 4.4 times its standard error). Content is preserved: cross-dial similarity (0.36) sits at or above the base model's own seed-to-seed floor (0.29), p = 0.002. The honest residual is magnitude, not existence: paired dz = 0.64 (bootstrap 95% CI 0.46 to 1.20) still straddles the 0.8 convention, so a clean magnitude headline wants a larger base; the controllability itself is established.
The GRPO arm
Before the boundary method broke through, GRPO trained a LoRA directly on the computable reward (no labels). The reward curve separates cleanly by target level: high-burstiness prompts hold near -0.57 to -1.08 across 110 steps while low-burstiness prompts sit around -4.7, so the reward distinguishes the two regimes. It did not, on its own, yield a reliable var(L) knob at distilgpt2 scale, the same noise wall the negative result documents. The boundary-steered dial is what finally moved the metric above noise while holding coherence.
Prompt-ceiling benchmark (Q1)
How close can prompting alone get to human rhythm?
Each model writes a news article for 200 Wikinews headlines (2016-2021, paired with the human article) under 5 non-overlapping prompting modes, 3 seeds each. Scoring seed-averages per headline first, then runs one paired Wilcoxon per mode and metric over the 200 matched pairs. Cells show the median Δ = human − model (matched rank-biserial in parentheses); positive means the model sits below the human article on that dimension, the flatness direction. *p<0.05, **p<0.01 (nominal, two-sided).
DeepSeek V4 Pro
| Mode | Bn | Gini | Seg. entropy | Memory M |
|---|---|---|---|---|
| Control | +0.035 (0.34)** | +0.018 (0.35)** | -0.050 (-0.12) | -0.025 (-0.07) |
| Instruction | +0.000 (0.05) | -0.005 (-0.02) | -0.145 (-0.47)** | +0.006 (0.07) |
| Few-shot | +0.029 (0.24)** | +0.014 (0.20)* | -0.068 (-0.26)** | +0.012 (0.01) |
| Plan | +0.013 (0.15) | +0.003 (0.08) | -0.117 (-0.36)** | -0.003 (0.03) |
| Self-refine | -0.018 (-0.10) | -0.010 (-0.11) | -0.139 (-0.49)** | +0.027 (0.14) |
Gemini 3.1 Pro
| Mode | Bn | Gini | Seg. entropy | Memory M |
|---|---|---|---|---|
| Control | +0.128 (0.85)** | +0.056 (0.81)** | -0.062 (-0.19)* | +0.043 (0.18)* |
| Instruction | +0.071 (0.56)** | +0.026 (0.47)** | -0.053 (-0.16) | +0.109 (0.38)** |
| Few-shot | +0.169 (0.94)** | +0.080 (0.93)** | +0.117 (0.39)** | +0.051 (0.24)** |
| Plan | +0.070 (0.58)** | +0.028 (0.48)** | -0.026 (-0.06) | +0.114 (0.47)** |
| Self-refine | +0.039 (0.27)** | +0.013 (0.21)** | -0.051 (-0.12) | +0.047 (0.12) |
Gemma 3 27B
| Mode | Bn | Gini | Seg. entropy | Memory M |
|---|---|---|---|---|
| Control | +0.069 (0.58)** | +0.037 (0.58)** | +0.063 (0.25)** | +0.023 (0.11) |
| Instruction | +0.049 (0.44)** | +0.025 (0.41)** | +0.086 (0.36)** | +0.040 (0.10) |
| Few-shot | +0.063 (0.65)** | +0.033 (0.63)** | +0.145 (0.48)** | +0.005 (0.02) |
| Plan | +0.061 (0.57)** | +0.033 (0.55)** | +0.129 (0.45)** | +0.003 (0.03) |
| Self-refine | +0.029 (0.30)** | +0.015 (0.29)** | +0.094 (0.32)** | +0.048 (0.21)** |
GLM-5.2
| Mode | Bn | Gini | Seg. entropy | Memory M |
|---|---|---|---|---|
| Control | +0.126 (0.85)** | +0.058 (0.82)** | +0.068 (0.25)** | +0.053 (0.19)* |
| Instruction | +0.061 (0.57)** | +0.029 (0.51)** | +0.064 (0.26)** | +0.076 (0.24)** |
| Few-shot | +0.148 (0.94)** | +0.068 (0.92)** | +0.221 (0.74)** | +0.022 (0.07) |
| Plan | +0.080 (0.65)** | +0.034 (0.59)** | +0.091 (0.32)** | +0.034 (0.20)* |
| Self-refine | +0.011 (0.13) | +0.002 (0.09) | +0.050 (0.33)** | -0.019 (-0.02) |
Grok 4.3
| Mode | Bn | Gini | Seg. entropy | Memory M |
|---|---|---|---|---|
| Control | +0.027 (0.27)** | +0.010 (0.23)** | -0.046 (-0.16) | -0.004 (0.01) |
| Instruction | +0.050 (0.42)** | +0.026 (0.39)** | +0.062 (0.19)* | +0.028 (0.10) |
| Few-shot | +0.108 (0.76)** | +0.052 (0.74)** | +0.180 (0.56)** | +0.001 (-0.05) |
| Plan | +0.061 (0.56)** | +0.026 (0.50)** | +0.136 (0.35)** | +0.023 (0.03) |
| Self-refine | +0.019 (0.19)* | +0.009 (0.14) | +0.055 (0.21)** | +0.028 (0.14) |
Grok 4.5
| Mode | Bn | Gini | Seg. entropy | Memory M |
|---|---|---|---|---|
| Control | -0.009 (-0.14) | -0.011 (-0.17)* | -0.249 (-0.74)** | -0.023 (-0.07) |
| Instruction | -0.025 (-0.20)* | -0.014 (-0.20)* | -0.099 (-0.40)** | +0.048 (0.22)** |
| Few-shot | +0.077 (0.69)** | +0.039 (0.68)** | +0.089 (0.28)** | -0.002 (-0.05) |
| Plan | -0.009 (-0.02) | -0.005 (-0.05) | -0.062 (-0.22)** | +0.021 (0.06) |
| Self-refine | -0.059 (-0.41)** | -0.033 (-0.44)** | -0.133 (-0.44)** | +0.071 (0.31)** |
Kimi K2.6
| Mode | Bn | Gini | Seg. entropy | Memory M |
|---|---|---|---|---|
| Control | +0.095 (0.69)** | +0.043 (0.63)** | -0.048 (-0.16)* | -0.025 (-0.09) |
| Instruction | -0.006 (-0.00) | -0.014 (-0.14) | -0.024 (-0.11) | -0.014 (0.08) |
| Few-shot | +0.123 (0.83)** | +0.063 (0.79)** | +0.152 (0.55)** | -0.073 (-0.26)** |
| Plan | +0.026 (0.24)** | +0.006 (0.09) | +0.032 (0.05) | +0.055 (0.17)* |
| Self-refine | -0.070 (-0.57)** | -0.040 (-0.58)** | +0.061 (0.16) | -0.052 (-0.23)** |
Claude Opus 4.8
| Mode | Bn | Gini | Seg. entropy | Memory M |
|---|---|---|---|---|
| Control | +0.119 (0.80)** | +0.054 (0.77)** | +0.116 (0.39)** | -0.069 (-0.24)** |
| Instruction | -0.037 (-0.25)** | -0.031 (-0.36)** | +0.032 (0.09) | -0.117 (-0.38)** |
| Few-shot | +0.039 (0.34)** | +0.012 (0.26)** | +0.253 (0.75)** | -0.045 (-0.25)** |
| Plan | -0.000 (0.01) | -0.004 (-0.10) | +0.108 (0.34)** | -0.059 (-0.25)** |
| Self-refine | -0.065 (-0.50)** | -0.043 (-0.59)** | -0.005 (0.01) | +0.039 (0.18)* |
Qwen3.6-27B
| Mode | Bn | Gini | Seg. entropy | Memory M |
|---|---|---|---|---|
| Control | +0.097 (0.67)** | +0.046 (0.67)** | -0.001 (0.10) | -0.007 (-0.05) |
| Instruction | +0.028 (0.34)** | +0.010 (0.28)** | -0.018 (-0.04) | -0.052 (-0.16)* |
| Few-shot | +0.119 (0.80)** | +0.060 (0.79)** | +0.127 (0.47)** | -0.071 (-0.26)** |
| Plan | +0.045 (0.41)** | +0.017 (0.35)** | +0.057 (0.16) | -0.028 (-0.07) |
| Self-refine | +0.039 (0.26)** | +0.021 (0.24)** | +0.029 (0.10) | +0.028 (0.09) |
Cross-model: the instruction mode head-to-head
Matched rank-biserial per dimension under a plain instruction, one row per generator family. Positive keeps the model below the human article; closer to 0 is closer to human rhythm.
| Model | Bn | Gini | Seg. entropy | Memory M |
|---|---|---|---|---|
| DeepSeek V4 Pro | 0.05 | -0.02 | -0.47 | 0.07 |
| Gemini 3.1 Pro | 0.56 | 0.47 | -0.16 | 0.38 |
| Gemma 3 27B | 0.44 | 0.41 | 0.36 | 0.10 |
| GLM-5.2 | 0.57 | 0.51 | 0.26 | 0.24 |
| Grok 4.3 | 0.42 | 0.39 | 0.19 | 0.10 |
| Grok 4.5 | -0.20 | -0.20 | -0.40 | 0.22 |
| Kimi K2.6 | -0.00 | -0.14 | -0.11 | 0.08 |
| Claude Opus 4.8 | -0.25 | -0.36 | 0.09 | -0.38 |
| Qwen3.6-27B | 0.34 | 0.28 | -0.04 | -0.16 |
This section regenerates from experiments/results/benchmark/<model>/seed_avg_paired.json, the same single source of truth behind the paper's Table 1: promoting a new model's results updates this page and the paper with zero hand edits. Raw generations and scored evaluations are archived per model in the HF datasets.
Reproducibility
Every figure regenerates from one source of truth.
The numbers and figures on this page are not hand-entered. They regenerate from experiments/results/figures_data.json, the single source of truth for the results, via make figures. The metric itself is pure standard library with unit tests, and only the true surprisal and the generation loop need the heavier dependencies.
Explore the live corpus dashboard or read the full pipeline and harness in the repository. The controllability verdict, the lambda sweep, the GRPO curve, and the ablation each have their own JSON record under experiments/results/, so any number above is traceable to a file.