Failure Register
Failure Register
Failure Register
ZeroModel keeps negative evidence visible because a method that produces a strong headline metric can still fail the property the system actually needs.
This page summarizes measured approaches that are unsupported or refuted under their recorded conditions. It does not generalize those failures beyond the stated fixtures, representations, calibration methods, models, or operating points.
Normalized pixels
Measured / unsupported
The corrected bounded arcade evaluation preserved strong raw ranking signal:
- raw top-1 exact-row accuracy: 1,008 / 1,344
- raw top-1 action accuracy: 1,302 / 1,344
- final benign coverage: 0 / 1,344
- false rejection rate: 1,344 / 1,344
The important result is the gap: action correctness substantially overstated state correctness, and the rejection calibration did not identify a useful governed operating region.
Frozen DINOv2 global retrieval
Measured / refuted within stated conditions
The tested global DINOv2 CLS systems did not materially improve governed policy addressing over normalized pixels under the recorded fixture and preprocessing. One measured system was roughly 32 exact-row percentage points below normalized pixels while recording 73.79% FAR and 292 conflicting-action errors.
This does not establish that learned or local representations cannot work.
Ridge linear probe
Measured / refuted within stated conditions
The rejection-equipped ridge probe recorded:
- 59.45% benign action accuracy
- 28.79% exact benign-row accuracy
- 100% FAR
- 521 conflicting-action errors
The recorded operating point did not justify promotion.
Registration + locality
Measured / unsupported
Bounded registration repaired part of the raw translation/locality ranking failure:
- raw top-1 exact row: 1,176 / 1,344
- raw top-1 action: 1,323 / 1,344
But the selected governed operating point still produced 0 / 1,344 final benign coverage.
Improved ranking was not treated as proof of a usable reader.
Frame-local discriminative evidence
Measured / refuted within stated conditions
The committed architecture-selection evidence records:
no_safe_architecture
no_feasible_operating_point
The benchmark machinery exists. The tested architecture family did not earn promotion.
Controlled PNG interventions against a fixed local provider
Measured / unsupported
Six identified representation runs held the provider configuration, prompt digest, parser, compiled policy, fixture identity, and inference settings fixed while changing only the recorded PNG representation recipe.
The tested implicit interventions did not improve the targeted behaviour:
cooldown-shape-v1: no targeted cooldown improvementcooldown-dual-v1: no targeted cooldown improvementcooldown-redundant-v1: action-changing errors increased from 1 to 2lane-enhanced-v1: action-changing errors increased from 1 to 2 and produced one rejected malformed response- labelled control: 8 / 8 exact
No implicit intervention passed the predefined smoke promotion gate, so no 112-state canonical promotion run followed.
Why publish these?
Because “the system chose the correct action” is not enough evidence that it identified the state correctly.
A method can look excellent under action accuracy while remaining unsafe or uninformative under exact-state correctness, rejection behaviour, coverage, or conflicting-action error. ZeroModel’s evidence model keeps those distinctions explicit.
Explore the claims registry โ ยท Reproduce current examples โ