OPERATIONS

Benchmarks

The fair rendered-image spatial benchmark: method, all four scoring conditions, confidence intervals, the per-capability audit table, and validity notes. Rowe's own score is a self-check and is labelled as one.

Run of 2026-09-02

What this page is

Rowe returns auditable metric geometry with error bars. When we benchmark, we publish the method, every scoring condition, a confidence interval, and the validity notes, and we cite the artifact the numbers came from. Rowe's own score on the same cases is a self-check of a deterministic geometry pipeline against its own calculator, and is labelled as one. None of it is evidence that Rowe out-reasons any other model.

This is the first completed run of the rendered-image benchmark we previously listed as in progress: 30 held-out cases, 10 capabilities with 3 cases each, 1 competitor model, 1 seed, 1 run, CPU only. It is published so the numbers can be audited, not so they can be cited as a ranking. Read the validity notes before quoting any figure.

Headline

Competitor: hunarbatra/SpatialThinker-3B

On the 7 capabilities that can carry a cross-model reading (classify, compare, fit, issue_localization, placement, predict, stability), the competitor scored 10 / 21 under the fair prompt and 10 / 21 under the prior JSON prompt (95% Wilson interval on the fair figure: 0.28–0.68). Across all 30 cases it scored 16 / 30 (53.3%; 95% Wilson interval 0.36–0.70) on the fair condition, against 13 / 30 on the prior setup. The net difference of +3 is entirely the 3 excluded, vocabulary-bound capabilities (retrieve, visual_grounding, workflow_feedback); on the citable set the paired change is zero.

Paired by case between the prior and fair prompts: 11 correct in both, 12 wrong in both, 5 wrong then correct (heldout.placement.016, heldout.stability.015, heldout.visual_grounding.008, heldout.visual_grounding.018, heldout.visual_grounding.028), 2 correct then wrong (heldout.issue_localization.003, heldout.placement.026).

Rowe, kept separate

Self-check, not a comparison

Rowe on its normal structured path: 30 / 30 correct, 0 wrong, 0 unparseable.

answer path learned_spatial_capability_head; checkpoint E:\Four-Echelon-LSM-eval\artifacts\checkpoints\active\rowe_spatial_capability_head.json

Rowe answered on its normal structured path: a small learned-threshold checkpoint (six scalar thresholds plus keyword weights) driving a deterministic geometry pipeline over the structured scene. The gold labels come from an equivalent deterministic geometry calculator reading the same structured input. A calculator agreeing with a calculator shows the pipeline is internally consistent on these cases. It is not a model result, it is not comparable to any competitor cell, and it is not evidence that Rowe out-reasons any external model.

This row is not in the competitor tables below on purpose. Where it appears in the per-capability table it is labelled self-check.

Competitor

All four scoring conditions

The same 30 cases were scored four ways so a reader can separate prompt effects from parser effects. Intervals are 95% Wilson score intervals on the raw count.

ConditionPromptScoringCorrectWrongUnparseableAccuracy95% interval
A. Prior prompt, strict JSON
A_unfair_prompt_strict_json
priorstrict JSON131700.43330.27–0.61
C. Prior prompt, tolerant extraction
C_unfair_prompt_tolerant
priortolerant extraction131700.43330.27–0.61
D. Fair prompt, strict JSON
D_fair_prompt_strict_json
fairstrict JSON00300.00000.00–0.11
B. Fair prompt, tolerant extraction
B_fair_prompt_tolerant
fairtolerant extraction161400.53330.36–0.70
A. Prior prompt, strict JSON — legend: repo case_prompt, JSON text scene, no image, strict JSON parse (prior setup)
The prior setup: the repo's case prompt with the scene as JSON text and a schema to fill, no image. Every output parsed as JSON, so tolerant re-scoring had nothing to recover.
C. Prior prompt, tolerant extraction — legend: same raw outputs as A, re-scored with format-tolerant extraction
Identical to A. Nothing was truncated in this run, so there is no output-format effect to correct for. The July format delta was truncation recovery, not a format effect.
D. Fair prompt, strict JSON — legend: rendered image + natural-language question, strict JSON parse
The fair prompt asks a natural-language question without requesting a schema, and the model answered in prose. Strict JSON parsing scores prose as unparseable on every case. Every one of these rows is a complete, non-empty prose answer; this row measures the parser, not the model.
B. Fair prompt, tolerant extraction — legend: rendered image + natural-language question + tolerant extraction (fair)
The fair condition: rendered image, task-native question, tolerant extraction. The net gain over A is exactly the three visual_grounding flips, a vocabulary-bound capability; on the citable capabilities the gains and losses cancel.

Audit table, not a ranking

Per capability

3 cases per capability. These rows exist so the aggregates above can be checked, and for no other purpose: at this sample size no per-capability figure is statistically meaningful, and the rows marked excluded are vocabulary tasks that say nothing about spatial reasoning. The Rowe column is the self-check described above.

CapabilityACDBRowe (self-check)Cross-model readingNote
classify3 / 33 / 30 / 33 / 33 / 3citable
compare0 / 30 / 30 / 30 / 33 / 3citableA genuine disagreement: in prose the model reported added or removed objects the scene does not have.
fit2 / 32 / 30 / 32 / 33 / 3citableYes/no question; a model answering at chance scores about half.
issue_localization3 / 33 / 30 / 32 / 33 / 3citableLost one case on the image prompt: it named the right pair but called the issue 'other' instead of a collision.
placement2 / 32 / 30 / 32 / 33 / 3citableYes/no question; gained one case and lost another between the two prompts.
predict0 / 30 / 30 / 30 / 33 / 3citableExtractor limitation under review: the freeform extractor recognises 'downward' but not 'downwards', which all three fair answers used, so the stated direction was not applied. Not a clean model miss. The fix changes scores and was deliberately left for a scored re-run.
stability0 / 30 / 30 / 31 / 33 / 3citableThe fair answers were self-contradictory ('No. The plate supports the block.'); the extractor takes the leading 'No'. The one correct case is a 'No' on a scene whose gold is unstable.
retrieve3 / 33 / 30 / 33 / 33 / 3excludedVocabulary-bound. The question lists the candidate asset ids and their tags, so this is keyword matching, not spatial reasoning. Excluded from cross-model conclusions.
visual_grounding0 / 30 / 30 / 33 / 33 / 3excludedVocabulary-bound. The fair question says 'bounding-box overlay', which the extractor maps straight to the expected overlay type; the JSON prompt did not hand over that phrase. These three flips are the entire net difference between the prior and fair prompts. Excluded from cross-model conclusions.
workflow_feedback0 / 30 / 30 / 30 / 33 / 3excludedScored against Rowe's internal replay-tag vocabulary. The three cases have no geometry to render, so the fair pass ran text-only. Excluded from cross-model conclusions.

How the run was done

Method

spatial_heldout_cases(30): deterministic held-out cases heldout.<capability>.000-029, three per capability across ten capabilities. Held out from Rowe's threshold fitting; the renders and the questions are ours.

27 of the 30 scenes were rendered to images; 3 cases have no geometry (heldout.workflow_feedback.009, heldout.workflow_feedback.019, heldout.workflow_feedback.029) and ran text-only in the fair pass. Every cell was scored by lsm.competitor_benchmark.evaluate_model_output.

Competitor setup

Weights
bfloat16, device_map=none, vision tower upcast to float32 (--fp32-vision), local files only
Decoding
Greedy (do_sample=False), one seed
Generation budget
96 new tokens per case; no output hit the cap (longest output 319 characters)
Wall-clock budget
3600 s per case; longest observed inference 562.7 s
Model load
65.6 s

Host

Machine
Dell laptop, Intel i5-8300H, 15.9 GB RAM
Accelerator
None. A GTX 1050 4 GB is present but unused; torch.cuda.is_available() was false. CPU only.
Stack
Python 3.11.9, torch 2.12.0+cpu, transformers 5.8.1, process at Idle priority

Timing

Mean per inference
3.9 min (fair pass 5.0 min, prior pass 2.7 min)
Total inference
231.4 min for 60 generations, plus model load
Run wall clock
01:57 to 05:50 Dell local, 2026-09-02

Read before citing anything above

Validity notes

  1. Three cases per capability. Thirty cases is three per capability. No per-capability number here carries an error bar worth quoting; the per-capability table is shown so the aggregate can be audited, not so any row can be cited. A perfect 3 of 3 has a 95% lower bound near 0.44, and a 0 of 3 an upper bound near 0.56.
  2. One competitor, one seed, one run. One 3B open-weight model, greedy decoding, one pass per condition. Nothing here generalises to 'external spatial models', and nothing here is a ranking.
  3. Three capabilities are decided by vocabulary. retrieve, visual_grounding, and workflow_feedback are scored against vocabulary the question or Rowe's internal schema supplies, so they say nothing about spatial reasoning and are excluded from any cross-model reading. The competitor headline on this page is therefore the seven-capability subset, and the competitor scored the same on it under the prior prompt and the fair prompt.
  4. Rowe's score is a self-check. Rowe's perfect score is a deterministic pipeline agreeing with the deterministic calculator that produced the gold labels. It is presented separately, labelled as a self-check, and must not be read as a lead over the competitor. The harness's own 'residual Rowe lead' line is arithmetic over that self-check and is not reproduced here.
  5. CPU inference on a laptop. Competitor inference ran in bfloat16 with an fp32 vision tower on a CPU at Idle priority. Timing figures are for this machine only, and numerics can differ from a GPU run; whether that moves any answer is unmeasured.
  6. The renders and the prompts are ours. The competitor answered Four Echelon's synthetic wireframe renders and Four Echelon's phrasing, scored by Four Echelon's extractor. A different renderer, phrasing, or extractor could move any cell, and one extractor limitation (predict) is already known.
  7. The earlier twelve-case attempt is invalid. An attempt on 2026-07-31 ran twelve cases under a 180 s per-case timeout that expired during CPU image prefill, truncating every rendered-image output to one token. It is recorded as defective and superseded by this run. Nothing from it is cited here.

Prior attempt, 2026-07-31: invalid (12 cases)

The first attempt ran twelve cases with the harness default of 180 s per case, which became generate(max_time=180). CPU image prefill alone took 182 to 493 s, so every rendered-image output was cut to a single token after the budget expired. Its image-condition numbers measured a wall-clock budget, not the model. That attempt is recorded as defective and is superseded by this run; do not cite it.

Where the numbers come from

Source

Report generated 2026-09-02T09:50:01Z. Artifacts live under E:\Four-Echelon-LSM-eval\artifacts\reports\fair_img_v2\ on the run host; the write-up is Four-Echelon-LSM docs/rowe-fair-benchmark-results-2026-09-02.md (2026-09-02), harness scripts/run_fair_spatial_benchmark.py.

fair_benchmark_report.json
Full report: per-condition rows with actual, expected, and extraction; Rowe rows; conditions legend; unrenderable cases. Every number on this page is read from it.
fair_benchmark_report.txt
The printed summary the harness wrote.
spatialthinker_raw.jsonl
60 raw generation records (30 cases x prior + fair): case id, condition, had_image, prompt_chars, seconds, text.
renders/
27 rendered scene PNGs; 3 cases have no geometry.
run-20260902.log
The single run log for the whole run.
python scripts/run_fair_spatial_benchmark.py --case-count 30 \
  --competitor spatialthinker=hunarbatra/SpatialThinker-3B \
  --fp32-vision --local-files-only --passes unfair,fair \
  --case-timeout-seconds 3600 --max-new-tokens 96 \
  --out-dir artifacts/reports/fair_img_v2

What we do and do not claim is listed in the evidence pack; the fixed-fixture smoke is the CAD eval sheet.