InconclusiveQualitativeReview —
openai/gpt-5.6-luna-pro
f4ac54dbc3
The exact Pile validation artifact, sequence treatment, and token-weighted perplexity aggregation used for Table 10 were not source-resolved; the published Pile value was not copied. · The paper adds OpenBookQA but does not define whether its ordinary or sequence-length-normalized accuracy is intended. Both evaluated alternatives, 19.6% and 30.8%, are retained without selecting one; the corresponding equal-weight averages are 47.0531% and 48.6531%. · The declared paper Markdown path was absent, so the fixed PDF was used for the Table 10 source text and row. · Unresolved metric interpretation: m-t10-m370-pile: The paper identifies this as perplexity on the Pile validation split with the GPT-NeoX tokenizer, but the released checkpoint run did not include that split. The exact validation artifact, sequence treatment, and token-weighted aggregation are not established by the available repository/evaluator inputs, so no Pile scalar is selected.; m-t10-m370-openbook: The current paper adds OpenBookQA to the inherited Mamba task set and labels the column generically as acc, but it does not state whether ordinary or sequence-length-normalized multiple-choice accuracy was used. The lm-eval 0.4.2 task emits both `acc,none` and `acc_norm,none`; the original Mamba metric definition excludes OpenBookQA, so it cannot resolve this added task. Both candidates are retained and neither is selected by numeric proximity to the published row.; m-t10-m370-avg: The source Table 10 Average is an equal-weight arithmetic average of the seven task columns, but OpenBookQA's ordinary-versus-normalized field is unresolved. The average therefore has two retained alternatives computed from the six source-resolved task fields plus OpenBookQA `acc,none` or `acc_norm,none`; no alternative is selected by proximity to the reported average. · Unresolved measurements: m-t10-m370-pile: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; m-t10-m370-openbook: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; m-t10-m370-avg: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值 · These seven measurements do not cover the full Table 10 row: Pile perplexity, OpenBookQA accuracy, and the seven-task average remain unresolved. · The declared workload contains 29 complete model rows, while the captured raw execution demonstrates only the Mamba-370M row. · The exact source dataset versions, reference execution conditions, and checkpoint-byte equivalence were not independently verified. · The declared paper Markdown source was absent; the fixed PDF was used for the reported Table 10 values. · The current execution supports approximate reproduction of the seven supplied Mamba-370M measurements: LAMBADA perplexity and six source-resolved task accuracies closely match Table 10 at displayed precision. However, those measurements do not establish the claim that the full Table 10 row was reproduced, because the Pile perplexity, OpenBookQA accuracy, and dependent average were not resolved, and the declared workload covered additional model rows beyond Mamba-370M. Some original claim measurements remain without evidence; available measurements are assessed individually.