Failure Register

Failure Register

Failure Register

ZeroModel keeps negative evidence visible because a method that produces a strong headline metric can still fail the property the system actually needs.

This page summarizes measured approaches that are unsupported or refuted under their recorded conditions. It does not generalize those failures beyond the stated fixtures, representations, calibration methods, models, or operating points.

Normalized pixels

Measured / unsupported

The corrected bounded arcade evaluation preserved strong raw ranking signal:

The important result is the gap: action correctness substantially overstated state correctness, and the rejection calibration did not identify a useful governed operating region.

Frozen DINOv2 global retrieval

Measured / refuted within stated conditions

The tested global DINOv2 CLS systems did not materially improve governed policy addressing over normalized pixels under the recorded fixture and preprocessing. One measured system was roughly 32 exact-row percentage points below normalized pixels while recording 73.79% FAR and 292 conflicting-action errors.

This does not establish that learned or local representations cannot work.

Ridge linear probe

Measured / refuted within stated conditions

The rejection-equipped ridge probe recorded:

The recorded operating point did not justify promotion.

Registration + locality

Measured / unsupported

Bounded registration repaired part of the raw translation/locality ranking failure:

But the selected governed operating point still produced 0 / 1,344 final benign coverage.

Improved ranking was not treated as proof of a usable reader.

Frame-local discriminative evidence

Measured / refuted within stated conditions

The committed architecture-selection evidence records:

no_safe_architecture
no_feasible_operating_point

The benchmark machinery exists. The tested architecture family did not earn promotion.

Controlled PNG interventions against a fixed local provider

Measured / unsupported

Six identified representation runs held the provider configuration, prompt digest, parser, compiled policy, fixture identity, and inference settings fixed while changing only the recorded PNG representation recipe.

The tested implicit interventions did not improve the targeted behaviour:

No implicit intervention passed the predefined smoke promotion gate, so no 112-state canonical promotion run followed.

Why publish these?

Because “the system chose the correct action” is not enough evidence that it identified the state correctly.

A method can look excellent under action accuracy while remaining unsafe or uninformative under exact-state correctness, rejection behaviour, coverage, or conflicting-action error. ZeroModel’s evidence model keeps those distinctions explicit.

Explore the claims registry โ†’ ยท Reproduce current examples โ†’