The paper’s claims are available in Research claims.
0 / 6 claims verified
the rest still being verified
The ablation conditions require the unavailable implementation, training corpus, checkpoints, and full 1.5B training recipe. Reconstructing only one ablation would not establish the matched controlled comparison.
0/2 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model_variant: LoGo · ablation_group: reference · reported_scope: single reported evaluation · parameter_scale: 1.5B | 73.74 percentage_points | — | Not assessed |
ablation: without progressive masking · model_variant: LoGo ablation · parameter_scale: 1.5B | 71.59 percentage_points | — | Not assessed |
Reported
57.65 pp
Observed
—
The report-defining 32k checkpoint, LoGo implementation, and exact benchmark task configurations are unavailable; the fixed dataset registry does not identify these task assets.
0/1 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model_variant: LoGo · reported_scope: single reported evaluation · attention_budget: 0.5 · checkpoint_context: 32k | 57.65 percentage_points | — | Not assessed |
The defining 1.5B comparison requires the unavailable LoGo and baseline checkpoints, 100B-token plus two context-extension training stages on the paper's in-house corpus, and an exact lm-evaluation-harness configuration. An independent reconstruction would require material unspecified substitutions and cannot receive strict automatic status.
0/3 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model_variant: LoGo · reported_scope: single reported evaluation · parameter_count: 1.5B · attention_budget: 0.5 · context_extension: 32k and 128k stages · pretraining_tokens: 100B | 2.112 loss | — | Not assessed |
dataset: Lambada · model_variant: LoGo · parameter_count: 1.5B · attention_budget: 0.5 | 7.5 perplexity | — | Not assessed |
model_variant: LoGo · reported_scope: single reported evaluation · parameter_count: 1.5B · attention_budget: 0.5 | 56.3 percentage_points | — | Not assessed |
Strict reproduction requires the author implementation, the in-house pretraining corpus, exact checkpoints or full training, and the complete seed and evaluation configuration; no verified repository or those assets are present in the fixed inputs.
0/3 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
comparison: full-attention Transformer · model_variant: LoGo · reported_scope: single run · parameter_count: 3.3B · training_corpus: in-house corpus similar to Seed-OSS | 1.996 loss | — | Not assessed |
dataset: WikiText · model_variant: LoGo · reported_scope: single run · parameter_count: 3.3B | 14.785 perplexity | — | Not assessed |
model_variant: LoGo · reported_scope: single reported evaluation · parameter_count: 3.3B | 58.26 percentage_points | — | Not assessed |
Reported
1.99 ratio
Observed
—
This hardware-sensitive claim needs the missing Triton kernel and dense comparator implementation, exact GPU and software versions, and a complete timing protocol including warmup, timed iterations, synchronization, and input construction. The paper does not determine all of these conditions.
0/1 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
direction: higher is faster · batch_size: 1 · comparator: dense FlashAttention Triton baseline · head_dimension: 128 · implementation: LoGo query-sparse Triton kernel · attention_heads: 8 · sequence_length: 65536 · attention_budget: 0.5 | 1.99 ratio | — | Not assessed |
The paper's context-extended checkpoints and exact RULER data/configuration are not provided, and no verified author repository exists. A substitute public model or reduced RULER panel would not cover the reported scope.
0/2 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
input_lengths: ["8k","16k","32k"] · model_variant: LoGo · subtask_count: 8 · checkpoint_context: 32k | 83 percentage_points | — | Not assessed |
input_lengths: ["32k","64k"] · model_variant: LoGo · subtask_count: 8 · checkpoint_context: 128k | 65.4 percentage_points | — | Not assessed |