OPERATIONS
Benchmarks
The fair rendered-image spatial benchmark: method, all four scoring conditions, confidence intervals, the per-capability audit table, and validity notes. Rowe's own score is a self-check and is labelled as one.
Run of 2026-09-02
What this page is
Rowe returns auditable metric geometry with error bars. When we benchmark, we publish the method, every scoring condition, a confidence interval, and the validity notes, and we cite the artifact the numbers came from. Rowe's own score on the same cases is a self-check of a deterministic geometry pipeline against its own calculator, and is labelled as one. None of it is evidence that Rowe out-reasons any other model.
This is the first completed run of the rendered-image benchmark we previously listed as in progress: 30 held-out cases, 10 capabilities with 3 cases each, 1 competitor model, 1 seed, 1 run, CPU only. It is published so the numbers can be audited, not so they can be cited as a ranking. Read the validity notes before quoting any figure.
Headline
Competitor: hunarbatra/SpatialThinker-3B
On the 7 capabilities that can carry a cross-model reading (classify, compare, fit, issue_localization, placement, predict, stability), the competitor scored 10 / 21 under the fair prompt and 10 / 21 under the prior JSON prompt (95% Wilson interval on the fair figure: 0.28–0.68). Across all 30 cases it scored 16 / 30 (53.3%; 95% Wilson interval 0.36–0.70) on the fair condition, against 13 / 30 on the prior setup. The net difference of +3 is entirely the 3 excluded, vocabulary-bound capabilities (retrieve, visual_grounding, workflow_feedback); on the citable set the paired change is zero.
Paired by case between the prior and fair prompts: 11 correct in both, 12 wrong in both, 5 wrong then correct (heldout.placement.016, heldout.stability.015, heldout.visual_grounding.008, heldout.visual_grounding.018, heldout.visual_grounding.028), 2 correct then wrong (heldout.issue_localization.003, heldout.placement.026).
Rowe, kept separate
Self-check, not a comparison
Rowe on its normal structured path: 30 / 30 correct, 0 wrong, 0 unparseable.
answer path learned_spatial_capability_head; checkpoint E:\Four-Echelon-LSM-eval\artifacts\checkpoints\active\rowe_spatial_capability_head.json
Rowe answered on its normal structured path: a small learned-threshold checkpoint (six scalar thresholds plus keyword weights) driving a deterministic geometry pipeline over the structured scene. The gold labels come from an equivalent deterministic geometry calculator reading the same structured input. A calculator agreeing with a calculator shows the pipeline is internally consistent on these cases. It is not a model result, it is not comparable to any competitor cell, and it is not evidence that Rowe out-reasons any external model.
This row is not in the competitor tables below on purpose. Where it appears in the per-capability table it is labelled self-check.
Competitor
All four scoring conditions
The same 30 cases were scored four ways so a reader can separate prompt effects from parser effects. Intervals are 95% Wilson score intervals on the raw count.
| Condition | Prompt | Scoring | Correct | Wrong | Unparseable | Accuracy | 95% interval |
|---|---|---|---|---|---|---|---|
| A. Prior prompt, strict JSON A_unfair_prompt_strict_json | prior | strict JSON | 13 | 17 | 0 | 0.4333 | 0.27–0.61 |
| C. Prior prompt, tolerant extraction C_unfair_prompt_tolerant | prior | tolerant extraction | 13 | 17 | 0 | 0.4333 | 0.27–0.61 |
| D. Fair prompt, strict JSON D_fair_prompt_strict_json | fair | strict JSON | 0 | 0 | 30 | 0.0000 | 0.00–0.11 |
| B. Fair prompt, tolerant extraction B_fair_prompt_tolerant | fair | tolerant extraction | 16 | 14 | 0 | 0.5333 | 0.36–0.70 |
- A. Prior prompt, strict JSON — legend: repo case_prompt, JSON text scene, no image, strict JSON parse (prior setup)
- The prior setup: the repo's case prompt with the scene as JSON text and a schema to fill, no image. Every output parsed as JSON, so tolerant re-scoring had nothing to recover.
- C. Prior prompt, tolerant extraction — legend: same raw outputs as A, re-scored with format-tolerant extraction
- Identical to A. Nothing was truncated in this run, so there is no output-format effect to correct for. The July format delta was truncation recovery, not a format effect.
- D. Fair prompt, strict JSON — legend: rendered image + natural-language question, strict JSON parse
- The fair prompt asks a natural-language question without requesting a schema, and the model answered in prose. Strict JSON parsing scores prose as unparseable on every case. Every one of these rows is a complete, non-empty prose answer; this row measures the parser, not the model.
- B. Fair prompt, tolerant extraction — legend: rendered image + natural-language question + tolerant extraction (fair)
- The fair condition: rendered image, task-native question, tolerant extraction. The net gain over A is exactly the three visual_grounding flips, a vocabulary-bound capability; on the citable capabilities the gains and losses cancel.
Audit table, not a ranking
Per capability
3 cases per capability. These rows exist so the aggregates above can be checked, and for no other purpose: at this sample size no per-capability figure is statistically meaningful, and the rows marked excluded are vocabulary tasks that say nothing about spatial reasoning. The Rowe column is the self-check described above.
| Capability | A | C | D | B | Rowe (self-check) | Cross-model reading | Note |
|---|---|---|---|---|---|---|---|
| classify | 3 / 3 | 3 / 3 | 0 / 3 | 3 / 3 | 3 / 3 | citable | |
| compare | 0 / 3 | 0 / 3 | 0 / 3 | 0 / 3 | 3 / 3 | citable | A genuine disagreement: in prose the model reported added or removed objects the scene does not have. |
| fit | 2 / 3 | 2 / 3 | 0 / 3 | 2 / 3 | 3 / 3 | citable | Yes/no question; a model answering at chance scores about half. |
| issue_localization | 3 / 3 | 3 / 3 | 0 / 3 | 2 / 3 | 3 / 3 | citable | Lost one case on the image prompt: it named the right pair but called the issue 'other' instead of a collision. |
| placement | 2 / 3 | 2 / 3 | 0 / 3 | 2 / 3 | 3 / 3 | citable | Yes/no question; gained one case and lost another between the two prompts. |
| predict | 0 / 3 | 0 / 3 | 0 / 3 | 0 / 3 | 3 / 3 | citable | Extractor limitation under review: the freeform extractor recognises 'downward' but not 'downwards', which all three fair answers used, so the stated direction was not applied. Not a clean model miss. The fix changes scores and was deliberately left for a scored re-run. |
| stability | 0 / 3 | 0 / 3 | 0 / 3 | 1 / 3 | 3 / 3 | citable | The fair answers were self-contradictory ('No. The plate supports the block.'); the extractor takes the leading 'No'. The one correct case is a 'No' on a scene whose gold is unstable. |
| retrieve | 3 / 3 | 3 / 3 | 0 / 3 | 3 / 3 | 3 / 3 | excluded | Vocabulary-bound. The question lists the candidate asset ids and their tags, so this is keyword matching, not spatial reasoning. Excluded from cross-model conclusions. |
| visual_grounding | 0 / 3 | 0 / 3 | 0 / 3 | 3 / 3 | 3 / 3 | excluded | Vocabulary-bound. The fair question says 'bounding-box overlay', which the extractor maps straight to the expected overlay type; the JSON prompt did not hand over that phrase. These three flips are the entire net difference between the prior and fair prompts. Excluded from cross-model conclusions. |
| workflow_feedback | 0 / 3 | 0 / 3 | 0 / 3 | 0 / 3 | 3 / 3 | excluded | Scored against Rowe's internal replay-tag vocabulary. The three cases have no geometry to render, so the fair pass ran text-only. Excluded from cross-model conclusions. |
How the run was done
Method
spatial_heldout_cases(30): deterministic held-out cases heldout.<capability>.000-029, three per capability across ten capabilities. Held out from Rowe's threshold fitting; the renders and the questions are ours.
27 of the 30 scenes were rendered to images; 3 cases have no geometry (heldout.workflow_feedback.009, heldout.workflow_feedback.019, heldout.workflow_feedback.029) and ran text-only in the fair pass. Every cell was scored by lsm.competitor_benchmark.evaluate_model_output.
Competitor setup
- Weights
- bfloat16, device_map=none, vision tower upcast to float32 (--fp32-vision), local files only
- Decoding
- Greedy (do_sample=False), one seed
- Generation budget
- 96 new tokens per case; no output hit the cap (longest output 319 characters)
- Wall-clock budget
- 3600 s per case; longest observed inference 562.7 s
- Model load
- 65.6 s
Host
- Machine
- Dell laptop, Intel i5-8300H, 15.9 GB RAM
- Accelerator
- None. A GTX 1050 4 GB is present but unused; torch.cuda.is_available() was false. CPU only.
- Stack
- Python 3.11.9, torch 2.12.0+cpu, transformers 5.8.1, process at Idle priority
Timing
- Mean per inference
- 3.9 min (fair pass 5.0 min, prior pass 2.7 min)
- Total inference
- 231.4 min for 60 generations, plus model load
- Run wall clock
- 01:57 to 05:50 Dell local, 2026-09-02
Read before citing anything above
Validity notes
- Three cases per capability. Thirty cases is three per capability. No per-capability number here carries an error bar worth quoting; the per-capability table is shown so the aggregate can be audited, not so any row can be cited. A perfect 3 of 3 has a 95% lower bound near 0.44, and a 0 of 3 an upper bound near 0.56.
- One competitor, one seed, one run. One 3B open-weight model, greedy decoding, one pass per condition. Nothing here generalises to 'external spatial models', and nothing here is a ranking.
- Three capabilities are decided by vocabulary. retrieve, visual_grounding, and workflow_feedback are scored against vocabulary the question or Rowe's internal schema supplies, so they say nothing about spatial reasoning and are excluded from any cross-model reading. The competitor headline on this page is therefore the seven-capability subset, and the competitor scored the same on it under the prior prompt and the fair prompt.
- Rowe's score is a self-check. Rowe's perfect score is a deterministic pipeline agreeing with the deterministic calculator that produced the gold labels. It is presented separately, labelled as a self-check, and must not be read as a lead over the competitor. The harness's own 'residual Rowe lead' line is arithmetic over that self-check and is not reproduced here.
- CPU inference on a laptop. Competitor inference ran in bfloat16 with an fp32 vision tower on a CPU at Idle priority. Timing figures are for this machine only, and numerics can differ from a GPU run; whether that moves any answer is unmeasured.
- The renders and the prompts are ours. The competitor answered Four Echelon's synthetic wireframe renders and Four Echelon's phrasing, scored by Four Echelon's extractor. A different renderer, phrasing, or extractor could move any cell, and one extractor limitation (predict) is already known.
- The earlier twelve-case attempt is invalid. An attempt on 2026-07-31 ran twelve cases under a 180 s per-case timeout that expired during CPU image prefill, truncating every rendered-image output to one token. It is recorded as defective and superseded by this run. Nothing from it is cited here.
Prior attempt, 2026-07-31: invalid (12 cases)
The first attempt ran twelve cases with the harness default of 180 s per case, which became generate(max_time=180). CPU image prefill alone took 182 to 493 s, so every rendered-image output was cut to a single token after the budget expired. Its image-condition numbers measured a wall-clock budget, not the model. That attempt is recorded as defective and is superseded by this run; do not cite it.
Where the numbers come from
Source
Report generated 2026-09-02T09:50:01Z. Artifacts live under E:\Four-Echelon-LSM-eval\artifacts\reports\fair_img_v2\ on the run host; the write-up is Four-Echelon-LSM docs/rowe-fair-benchmark-results-2026-09-02.md (2026-09-02), harness scripts/run_fair_spatial_benchmark.py.
- fair_benchmark_report.json
- Full report: per-condition rows with actual, expected, and extraction; Rowe rows; conditions legend; unrenderable cases. Every number on this page is read from it.
- fair_benchmark_report.txt
- The printed summary the harness wrote.
- spatialthinker_raw.jsonl
- 60 raw generation records (30 cases x prior + fair): case id, condition, had_image, prompt_chars, seconds, text.
- renders/
- 27 rendered scene PNGs; 3 cases have no geometry.
- run-20260902.log
- The single run log for the whole run.
python scripts/run_fair_spatial_benchmark.py --case-count 30 \
--competitor spatialthinker=hunarbatra/SpatialThinker-3B \
--fp32-vision --local-files-only --passes unfair,fair \
--case-timeout-seconds 3600 --max-new-tokens 96 \
--out-dir artifacts/reports/fair_img_v2What we do and do not claim is listed in the evidence pack; the fixed-fixture smoke is the CAD eval sheet.