SigmaLens

Full method

Evidence and validation

One page, for a reader who wants to know what has actually been demonstrated before looking at anything else. It includes the experiments that failed and the two figures we withdrew, because a validation page that only carries what survived is a sales document. Every number below is served live from /claims, which is the same register the product itself renders. Nothing on this page is typed in by hand.

What is true of this system right now

Measured on this deployment, not asserted. Each line is either true of the records in the database or it is not, and the script that decides is tools/audit_claims.py, which ships with the product. None of it is a performance claim, and after the record below that is deliberate: what this product can be held to is how it behaves, not how it scores.

Measuring…

The claim register

Every performance claim this product has made, with its current status. A row with nothing in the last column may not be quoted anywhere, including by us.

Loading the register…

What we claim

One thing, and it is not a ranking claim and not a return claim:

An event reaches a book, or it does not, and the path is auditable

Given a portfolio, SigmaLens states which holdings a given event reaches, through which documented channel, at what remove, with the evidence and effective date attached, and it states explicitly which holdings it does not reach and which entities it cannot map at all. Every relationship in the graph carries a citation and a date. Nothing is inferred from a sector code.

That is a capability claim about what the system produces and whether the production is auditable. It is not a claim that acting on the output makes money, and we have looked for that five times and not found it.

What we have demonstrated

It declines, and it says why

Counted across every day of the frozen panel, not asserted. The rate and the reasons are in the register above. An entity with no evidenced economic channel is returned as unmapped rather than given a relationship, and an entity below the evidence minimum is returned as insufficient rather than ranked low.

The exposure graph is a documented artifact, not a heuristic

Each edge carries a source, a rationale, an effective date and a structural confidence. The whole schema is fingerprinted, and the fingerprint is compared at serve time, so a change to the graph cannot silently invalidate a published result again.

The score agrees with a blinded AI evaluator panel, and so does counting

Exploratory. On the 42 evaluation days, the engine's salience is rank correlated with blinded materiality at about +0.67. Counting events on the same rows gets about +0.66. The reading is that the score is not noise and is not better than the cheapest alternative, which is the same conclusion the ranking experiment reached by a different route.

What remains unproven

What failed

The ranking claim, withdrawn 2 September 2026

We published that the ordering beat a plain event-volume sort. It does not. The evaluation harness had been building its exposure graph without the entity map the product uses, so every published block measured an engine that was never shipped. Rebuilt on the shipped configuration across all three preregistered blocks with fresh gold panels and 45 blinded AI-evaluator runs (not human professionals), the result is −0.05 mean, Fisher p = 0.26, 20 of 42 days. On the 2023 to 2024 block the volume sort won outright.

The directional claim, withdrawn earlier

67.8% against a 54.3% base rate came from a hand-picked library of 83 events. Replayed point in time, with nobody choosing the events, the same pipeline scores 47.6% on the densest sample and 48.3% out of sample, and the calls lose money net of costs.

The sentiment materiality result, withdrawn 3 September 2026

We published that the severity score ranks by materiality at rho +0.48. Audited during this claims review: all 120 cases in that study carried a "headline" the system had written itself, of the form Commodity coverage (Crude Oil): 22 article(s), avg tone -1.5, and the evaluators were shown the source count and reliability tier that the severity score is computed from. The study could not separate our score from our own summary of its own inputs. Real published titles have been available since the feed fix; the study has not been re-run on them, so the figure is withdrawn rather than restated.

The crowding explanation, falsified

We predicted in advance that the ordering advantage would be largest on crowded days and preregistered a block of them. It came back negative. There is no known rule for when the ordering helps, and we stopped claiming one.

How we test ourselves

Current product capabilities

Current limitations

Full method, all three blocks, every robustness check →