Explore ArkGraph and select the steps to run.
The paper’s claims are available in Research claims.
0 / 19 claims verified
the rest still being verified
Experiment plan ready; no runs yet.
0/36 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
tier: 10M | 6 count | — | Not assessed |
tier: 10M | 6 count | — | Not assessed |
tier: 10M | 384 dimensions | — | Not assessed |
tier: 10M | 512 tokens | — | Not assessed |
tier: 10M | 64 sequences | — | Not assessed |
tier: 10M | 1 steps | — | Not assessed |
tier: 10M | 6,000 steps | — | Not assessed |
tier: 10M | 200 steps | — | Not assessed |
tier: 10M | 0.2 billions of tokens | — | Not assessed |
tier: 50M | 10 count | — | Not assessed |
tier: 50M | 10 count | — | Not assessed |
tier: 50M | 640 dimensions | — | Not assessed |
Experiment plan ready; no runs yet.
0/16 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
arm: norm gate · tier: 124M · override: trained gate · dose_epsilon: 0 | 3.3472 nats | — | Not assessed |
arm: norm gate · tier: 124M · override: trained gate · junk_type: structured · dose_epsilon: 0.2 | 3.3628 nats | — | Not assessed |
arm: norm gate · tier: 124M · override: row mean · dose_epsilon: 0 | 3.5376 nats | — | Not assessed |
arm: norm gate · tier: 124M · override: row mean · junk_type: structured · dose_epsilon: 0.2 | 3.6606 nats | — | Not assessed |
arm: norm gate · tier: 124M · override: shuffled · dose_epsilon: 0 | 3.6607 nats | — | Not assessed |
arm: norm gate · tier: 124M · override: shuffled · junk_type: structured · dose_epsilon: 0.2 | 3.8877 nats | — | Not assessed |
arm: norm gate · tier: 124M · override: suppressed reads reopened · dose_epsilon: 0 | 11.6894 nats | — | Not assessed |
arm: norm gate · tier: 124M · override: suppressed reads reopened · junk_type: structured · dose_epsilon: 0.2 | 11.8153 nats | — | Not assessed |
arm: combo · tier: 124M · override: trained gate · dose_epsilon: 0 | 3.3371 nats | — | Not assessed |
arm: combo · tier: 124M · override: trained gate · junk_type: structured · dose_epsilon: 0.2 | 3.3458 nats | — | Not assessed |
arm: combo · tier: 124M · override: row mean · dose_epsilon: 0 | 3.5034 nats | — | Not assessed |
arm: combo · tier: 124M · override: row mean · junk_type: structured · dose_epsilon: 0.2 | 3.5408 nats | — | Not assessed |
Experiment plan ready; no runs yet.
0/30 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
arm: baseline · seed: 0 · tier: 124M · cutoff: median | 0.564 fraction of attention-row mass | — | Not assessed |
arm: baseline · seed: 0 · tier: 124M · cutoff: q25 | 0.373 fraction of attention-row mass | — | Not assessed |
arm: baseline · seed: 0 · tier: 50M · cutoff: median | 0.554 fraction of attention-row mass | — | Not assessed |
arm: baseline · seed: 0 · tier: 50M · cutoff: q25 | 0.357 fraction of attention-row mass | — | Not assessed |
arm: sink logit · seed: 0 · tier: 124M · cutoff: median | 0.312 fraction of attention-row mass | — | Not assessed |
arm: sink logit · seed: 0 · tier: 124M · cutoff: q25 | 0.18 fraction of attention-row mass | — | Not assessed |
arm: sink logit · seed: 0 · tier: 124M · cutoff: not applicable | 0.352 fraction of attention-row mass | — | Not assessed |
arm: sink logit · seed: 0 · tier: 50M · cutoff: median | 0.338 fraction of attention-row mass | — | Not assessed |
arm: sink logit · seed: 0 · tier: 50M · cutoff: q25 | 0.196 fraction of attention-row mass | — | Not assessed |
arm: sink logit · seed: 0 · tier: 50M · cutoff: not applicable | 0.309 fraction of attention-row mass | — | Not assessed |
arm: norm gate · seed: 0 · tier: 124M · cutoff: median | 0.617 fraction of attention-row mass | — | Not assessed |
arm: norm gate · seed: 0 · tier: 124M · cutoff: q25 | 0.443 fraction of attention-row mass | — | Not assessed |
Experiment plan ready; no runs yet.
0/328 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
arm: baseline · seed: 0 · tier: FineWeb 10M · cutoff: q25 · junk_type: structured other context · dose_epsilon: 0.2 | 0.032 nats | — | Not assessed |
arm: baseline · seed: 0 · tier: FineWeb 10M · cutoff: q25 · junk_type: structured other context · dose_epsilon: 0.4 | 0.138 nats | — | Not assessed |
arm: baseline · seed: 0 · tier: FineWeb 10M · cutoff: q25 · junk_type: structured other context · dose_epsilon: 0.8 | 0.771 nats | — | Not assessed |
arm: baseline · seed: 0 · tier: FineWeb 10M · cutoff: q25 · junk_type: structured other context · dose_epsilon: 1.6 | 2.695 nats | — | Not assessed |
arm: sink logit · seed: 0 · tier: FineWeb 10M · cutoff: q25 · junk_type: structured other context · dose_epsilon: 0.2 | 0.009 nats | — | Not assessed |
arm: sink logit · seed: 0 · tier: FineWeb 10M · cutoff: q25 · junk_type: structured other context · dose_epsilon: 0.4 | 0.041 nats | — | Not assessed |
arm: sink logit · seed: 0 · tier: FineWeb 10M · cutoff: q25 · junk_type: structured other context · dose_epsilon: 0.8 | 0.178 nats | — | Not assessed |
arm: sink logit · seed: 0 · tier: FineWeb 10M · cutoff: q25 · junk_type: structured other context · dose_epsilon: 1.6 | 0.778 nats | — | Not assessed |
arm: norm gate · seed: 0 · tier: FineWeb 10M · cutoff: q25 · junk_type: structured other context · dose_epsilon: 0.2 | 0.027 nats | — | Not assessed |
arm: norm gate · seed: 0 · tier: FineWeb 10M · cutoff: q25 · junk_type: structured other context · dose_epsilon: 0.4 | 0.124 nats | — | Not assessed |
arm: norm gate · seed: 0 · tier: FineWeb 10M · cutoff: q25 · junk_type: structured other context · dose_epsilon: 0.8 | 0.877 nats | — | Not assessed |
arm: norm gate · seed: 0 · tier: FineWeb 10M · cutoff: q25 · junk_type: structured other context · dose_epsilon: 1.6 | 3.162 nats | — | Not assessed |
Experiment plan ready; no runs yet.
0/36 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
tier: 10M · gate_form: norm · comparison: baseline except filter increment, which is both minus sink | 0.0185 nats | — | Not assessed |
tier: 10M · gate_form: norm · comparison: baseline except filter increment, which is both minus sink | 0.0037 nats | — | Not assessed |
tier: 10M · gate_form: norm · comparison: baseline except filter increment, which is both minus sink | 0.0177 nats | — | Not assessed |
tier: 10M · gate_form: norm · comparison: baseline except filter increment, which is both minus sink | 0.0045 nats | — | Not assessed |
tier: 10M · gate_form: norm · comparison: baseline except filter increment, which is both minus sink | -0.0008 nats | — | Not assessed |
tier: 50M · gate_form: norm · comparison: baseline except filter increment, which is both minus sink | 0.0068 nats | — | Not assessed |
tier: 50M · gate_form: norm · comparison: baseline except filter increment, which is both minus sink | 0.0069 nats | — | Not assessed |
tier: 50M · gate_form: norm · comparison: baseline except filter increment, which is both minus sink | 0.0106 nats | — | Not assessed |
tier: 50M · gate_form: norm · comparison: baseline except filter increment, which is both minus sink | 0.003 nats | — | Not assessed |
tier: 50M · gate_form: norm · comparison: baseline except filter increment, which is both minus sink | 0.003 nats | — | Not assessed |
tier: 124M · gate_form: norm · comparison: baseline except filter increment, which is both minus sink | 0.0052 nats | — | Not assessed |
tier: 124M · gate_form: norm · comparison: baseline except filter increment, which is both minus sink | 0.0059 nats | — | Not assessed |
Experiment plan ready; no runs yet.
0/20 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
arm: baseline · tier: 124M · examples: 953 · target_distance_band: target absent | 0.028 fraction | — | Not assessed |
arm: baseline · tier: 124M · examples: 1508 · target_distance_band: distance <=32 | 0.253 fraction | — | Not assessed |
arm: baseline · tier: 124M · examples: 2660 · target_distance_band: distance 33-96 | 0.198 fraction | — | Not assessed |
arm: baseline · tier: 124M · examples: 32 · target_distance_band: distance >96 | 0.104 fraction | — | Not assessed |
arm: sink logit · tier: 124M · examples: 953 · target_distance_band: target absent | 0.03 fraction | — | Not assessed |
arm: sink logit · tier: 124M · examples: 1508 · target_distance_band: distance <=32 | 0.26 fraction | — | Not assessed |
arm: sink logit · tier: 124M · examples: 2660 · target_distance_band: distance 33-96 | 0.188 fraction | — | Not assessed |
arm: sink logit · tier: 124M · examples: 32 · target_distance_band: distance >96 | 0.115 fraction | — | Not assessed |
arm: norm gate · tier: 124M · examples: 953 · target_distance_band: target absent | 0.031 fraction | — | Not assessed |
arm: norm gate · tier: 124M · examples: 1508 · target_distance_band: distance <=32 | 0.26 fraction | — | Not assessed |
arm: norm gate · tier: 124M · examples: 2660 · target_distance_band: distance 33-96 | 0.198 fraction | — | Not assessed |
arm: norm gate · tier: 124M · examples: 32 · target_distance_band: distance >96 | 0.104 fraction | — | Not assessed |
Experiment plan ready; no runs yet.
0/30 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
arm: baseline · tier: 124M · zero_shot: true | 0.182 fraction | — | Not assessed |
arm: baseline · tier: 124M · zero_shot: true | 0.006 fraction | — | Not assessed |
arm: baseline · tier: 124M · zero_shot: true | 0.323 fraction | — | Not assessed |
arm: baseline · tier: 124M · zero_shot: true | 0.002 fraction | — | Not assessed |
arm: baseline · tier: 124M · zero_shot: true | 58.1 perplexity | — | Not assessed |
arm: baseline · tier: 124M · zero_shot: true | 0.5 perplexity | — | Not assessed |
arm: sink logit · tier: 124M · zero_shot: true | 0.179 fraction | — | Not assessed |
arm: sink logit · tier: 124M · zero_shot: true | 0.004 fraction | — | Not assessed |
arm: sink logit · tier: 124M · zero_shot: true | 0.325 fraction | — | Not assessed |
arm: sink logit · tier: 124M · zero_shot: true | 0.003 fraction | — | Not assessed |
arm: sink logit · tier: 124M · zero_shot: true | 57 perplexity | — | Not assessed |
arm: sink logit · tier: 124M · zero_shot: true | 0.5 perplexity | — | Not assessed |
Experiment plan ready; no runs yet.
0/70 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: gpt2 | 124 millions of parameters | — | Not assessed |
model: gpt2 | 12 count | — | Not assessed |
model: gpt2 | 12 count | — | Not assessed |
model: gpt2 | 0.364 fraction | — | Not assessed |
model: gpt2 | 0.144 ratio | — | Not assessed |
model: gpt2 | 1.03 ratio | — | Not assessed |
model: gpt2 | 0.258 fraction | — | Not assessed |
model: gpt2-medium | 355 millions of parameters | — | Not assessed |
model: gpt2-medium | 24 count | — | Not assessed |
model: gpt2-medium | 16 count | — | Not assessed |
model: gpt2-medium | 0.388 fraction | — | Not assessed |
model: gpt2-medium | 0.092 ratio | — | Not assessed |
Experiment plan ready; no runs yet.
0/69 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
arm: baseline · tier: 10M | 4.2119 nats | — | Not assessed |
arm: baseline · tier: 50M | 3.5399 nats | — | Not assessed |
arm: baseline · tier: 124M | 3.371 nats | — | Not assessed |
arm: baseline · tier: 350M | 3.0338 nats | — | Not assessed |
arm: off-by-one · tier: 10M · comparison: baseline | -0.016 nats | — | Not assessed |
arm: off-by-one · tier: 10M · comparison: baseline | 0.0027 nats | — | Not assessed |
arm: sink token · tier: 10M · comparison: baseline | -0.0025 nats | — | Not assessed |
arm: sink token · tier: 10M · comparison: baseline | 0.0041 nats | — | Not assessed |
arm: sink logit · tier: 10M · comparison: baseline | -0.0185 nats | — | Not assessed |
arm: sink logit · tier: 10M · comparison: baseline | 0.0033 nats | — | Not assessed |
arm: sink logit · tier: 50M · comparison: baseline | -0.0068 nats | — | Not assessed |
arm: sink logit · tier: 50M · comparison: baseline | 0.0003 nats | — | Not assessed |
Experiment plan ready; no runs yet.
0/9 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
arm: norm gate · seed: 0 · tier: 10M | 0.054 gate-value units | — | Not assessed |
arm: norm gate · seed: 0 · tier: 50M | 0.17 gate-value units | — | Not assessed |
arm: norm gate · seed: 0 · tier: 124M | 0.177 gate-value units | — | Not assessed |
arm: combo · seed: 0 · tier: 10M | 0.025 gate-value units | — | Not assessed |
arm: combo · seed: 0 · tier: 50M | 0.138 gate-value units | — | Not assessed |
arm: combo · seed: 0 · tier: 124M | 0.158 gate-value units | — | Not assessed |
arm: combo · seed: 0 · tier: 350M | 0.151 gate-value units | — | Not assessed |
arm: norm gate · seed: 0 · tier: 124M | 0.46 gate-value units | — | Not assessed |
arm: combo · seed: 0 · tier: 124M | 0.58 gate-value units | — | Not assessed |
Experiment plan ready; no runs yet.
0/7 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
arm: baseline · tier: 124M | 0.6 +excess kurtosis | — | Not assessed |
arm: sink logit · tier: 124M | 0.49 +excess kurtosis | — | Not assessed |
arm: norm gate · tier: 124M | 0.74 +excess kurtosis | — | Not assessed |
arm: combo · tier: 124M | 0.68 +excess kurtosis | — | Not assessed |
arm: projection gate · tier: 124M | 0.38 +excess kurtosis | — | Not assessed |
arm: combo2 · tier: 124M | 0.34 +excess kurtosis | — | Not assessed |
arm: combo · tier: 350M | 0.82 +excess kurtosis | — | Not assessed |
Experiment plan ready; no runs yet.
0/4 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
arm: norm gate · tier: 124M · tokens: newline, period, comma, semicolon, opening parenthesis, the/The/a/and | 0.36 gate-value units | — | Not assessed |
arm: norm gate · tier: 124M · tokens: newline, period, comma, semicolon, opening parenthesis, the/The/a/and | 0.43 gate-value units | — | Not assessed |
arm: norm gate · tier: 124M · tokens: which, not, can, will, more, you | 0.5 gate-value units | — | Not assessed |
arm: norm gate · tier: 124M · tokens: which, not, can, will, more, you | 0.56 gate-value units | — | Not assessed |
Experiment plan ready; no runs yet.
0/14 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
arm: baseline · tier: 124M | 0.06 fraction | — | Not assessed |
arm: baseline · tier: 124M · denominator: ordinary read norm | 0.42 ratio | — | Not assessed |
arm: sink logit · tier: 124M | 0.005 fraction | — | Not assessed |
arm: sink logit · tier: 124M · denominator: ordinary read norm | 0.99 ratio | — | Not assessed |
arm: norm gate · tier: 124M | 0.037 fraction | — | Not assessed |
arm: norm gate · tier: 124M · denominator: ordinary read norm | 0.65 ratio | — | Not assessed |
arm: baseline · tier: 350M | 0.11 fraction | — | Not assessed |
arm: baseline · tier: 350M · denominator: ordinary read norm | 0.36 ratio | — | Not assessed |
arm: sink logit · tier: 350M | 0.007 fraction | — | Not assessed |
arm: sink logit · tier: 350M · denominator: ordinary read norm | 0.98 ratio | — | Not assessed |
arm: sink logit · tier: 350M | 0.53 fraction | — | Not assessed |
arm: sink logit · tier: 10M | 0.13 fraction | — | Not assessed |
Experiment plan ready; no runs yet.
0/114 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
arm: baseline · tier: TinyStories 10M · statistic: seed 0 | 1.5881 nats | — | Not assessed |
arm: baseline · tier: TinyStories 10M · statistic: seed 1 | 1.587 nats | — | Not assessed |
arm: baseline · tier: TinyStories 10M · statistic: seed 2 | 1.5869 nats | — | Not assessed |
arm: baseline · tier: TinyStories 10M · statistic: mean | 1.5873 nats | — | Not assessed |
arm: off-by-one · tier: TinyStories 10M · statistic: seed 0 | 1.5852 nats | — | Not assessed |
arm: off-by-one · tier: TinyStories 10M · statistic: seed 1 | 1.5826 nats | — | Not assessed |
arm: off-by-one · tier: TinyStories 10M · statistic: seed 2 | 1.5835 nats | — | Not assessed |
arm: off-by-one · tier: TinyStories 10M · statistic: mean | 1.5838 nats | — | Not assessed |
arm: renorm twin · tier: TinyStories 10M · statistic: seed 0 | 1.588 nats | — | Not assessed |
arm: renorm twin · tier: TinyStories 10M · statistic: seed 1 | 1.5869 nats | — | Not assessed |
arm: renorm twin · tier: TinyStories 10M · statistic: seed 2 | 1.5876 nats | — | Not assessed |
arm: renorm twin · tier: TinyStories 10M · statistic: mean | 1.5875 nats | — | Not assessed |
The task is fully specified and cheap to rerun. The paper reports only qualitative Figure 3 relations, so raw curves and source-context assessment are retained instead of a fabricated numeric target.
The 10M high-dose magnitude failure is testable in the 10M battery. The projection-direction comparison is retained in 50M/124M packages but currently requires resource-inadmissible checkpoints; the inventory has no standalone numeric IDs for this qualitative relation.
This is protocol context rather than an independently reported result; it is enforced by every injection package.
All nine arms and the decisive 10M controls are reconstructed in one matched package; arm equations require execution-time semantic checks because no author code exists.
This is the paper's limitation statement, not an independently testable reported result. The catalog preserves rather than erases those limits.