Evidence and validation
One page, for a reader who wants to know what has actually been demonstrated before looking at anything else. It includes the experiments that failed and the two figures we withdrew, because a validation page that only carries what survived is a sales document. Every number below is served live from /claims, which is the same register the product itself renders. Nothing on this page is typed in by hand.
What is true of this system right now
Measured on this deployment, not asserted. Each line is either true of
the records in the database or it is not, and the script that decides is
tools/audit_claims.py, which ships with the product. None of
it is a performance claim, and after the record below that is deliberate:
what this product can be held to is how it behaves, not how it scores.
The claim register
Every performance claim this product has made, with its current status. A row with nothing in the last column may not be quoted anywhere, including by us.
What we claim
One thing, and it is not a ranking claim and not a return claim:
An event reaches a book, or it does not, and the path is auditable
Given a portfolio, SigmaLens states which holdings a given event reaches, through which documented channel, at what remove, with the evidence and effective date attached, and it states explicitly which holdings it does not reach and which entities it cannot map at all. Every relationship in the graph carries a citation and a date. Nothing is inferred from a sector code.
That is a capability claim about what the system produces and whether the production is auditable. It is not a claim that acting on the output makes money, and we have looked for that five times and not found it.
What we have demonstrated
It declines, and it says why
Counted across every day of the frozen panel, not asserted. The rate and the reasons are in the register above. An entity with no evidenced economic channel is returned as unmapped rather than given a relationship, and an entity below the evidence minimum is returned as insufficient rather than ranked low.
The exposure graph is a documented artifact, not a heuristic
Each edge carries a source, a rationale, an effective date and a structural confidence. The whole schema is fingerprinted, and the fingerprint is compared at serve time, so a change to the graph cannot silently invalidate a published result again.
The score agrees with a blinded AI evaluator panel, and so does counting
Exploratory. On the 42 evaluation days, the engine's salience is rank correlated with blinded materiality at about +0.67. Counting events on the same rows gets about +0.66. The reading is that the score is not noise and is not better than the cheapest alternative, which is the same conclusion the ranking experiment reached by a different route.
What remains unproven
- Any return effect. Five preregistered studies, including one over 1,064 days. None positive. Materiality is not predictability and we do not present it as one.
- That correct attribution is what helps a reader. Book-addressed attribution beats no attribution, but against a control that was confidently wrong in the reader's own book the difference was not detectable: readers audited and corrected it either way. Content and format are still confounded.
- That the stated confidence label means anything. It reads HIGH on almost every scored record. What actually carries uncertainty in this system is the decision to decline, not the label.
- That declining is the right call. We can show what the engine declines and why. We have never collected blinded materiality for the records it declined, so we cannot yet show that they were less material than the ones it kept. That panel is the next experiment.
What failed
The ranking claim, withdrawn 2 September 2026
We published that the ordering beat a plain event-volume sort. It does not. The evaluation harness had been building its exposure graph without the entity map the product uses, so every published block measured an engine that was never shipped. Rebuilt on the shipped configuration across all three preregistered blocks with fresh gold panels and 45 blinded AI-evaluator runs (not human professionals), the result is −0.05 mean, Fisher p = 0.26, 20 of 42 days. On the 2023 to 2024 block the volume sort won outright.
The directional claim, withdrawn earlier
67.8% against a 54.3% base rate came from a hand-picked library of 83 events. Replayed point in time, with nobody choosing the events, the same pipeline scores 47.6% on the densest sample and 48.3% out of sample, and the calls lose money net of costs.
The sentiment materiality result, withdrawn 3 September 2026
We published that the severity score ranks by materiality at rho +0.48. Audited during this claims review: all 120 cases in that study carried a "headline" the system had written itself, of the form Commodity coverage (Crude Oil): 22 article(s), avg tone -1.5, and the evaluators were shown the source count and reliability tier that the severity score is computed from. The study could not separate our score from our own summary of its own inputs. Real published titles have been available since the feed fix; the study has not been re-run on them, so the figure is withdrawn rather than restated.
The crowding explanation, falsified
We predicted in advance that the ordering advantage would be largest on crowded days and preregistered a block of them. It came back negative. There is no known rule for when the ordering helps, and we stopped claiming one.
How we test ourselves
- Preregistration before materials exist. Dates, arms, statistics and the bar are fixed in a file in the repository before anything is generated, and the file is committed first.
- Blinded evaluation with separate contexts. Five independent professional lenses, none told which system produced what they saw. What counts as material is scored by different evaluators who saw no ordering at all.
- The cheapest alternative is always the baseline. Not "nothing". The comparison is against the list a competent person could build in an afternoon, because that is the real alternative.
- A seed-free headline statistic. An exact two-sided sign-flip permutation test, with the bootstrap reported beside it and its stability checked across twenty seeds, because a bootstrap interval once crossed zero depending on which seed it was given.
- Production and evaluation parity. After the ranking retraction, the engine records a hash of its own configuration alongside each result and compares it at serve time.
- Failures are never deleted. Superseded blocks stay in the repository with their raw evaluator output and still reproduce their own numbers, so both records can be checked against each other.
Current product capabilities
- Ingests continuously. Public event and coverage data on a fixed schedule, with real published titles recovered from the source records.
- Resolves with citations. An event to entities, and an entity to instruments through a curated channel graph with citations, effective dates and explicit direct, indirect and unmapped classes.
- Answers on request, stores nothing. A portfolio question against a book supplied on the request, returning which holdings are reached, which are not, and why for both.
- Declines rather than guesses. Refuses to score what it cannot evidence, and returns the reason rather than a low number.
- Ships as an API. Sits inside an existing workflow rather than asking anyone to visit a website.
Current limitations
- Coverage is partial and stated. A minority of the entities we scan reach no instrument at all. They are reported as unmapped rather than mapped optimistically.
- The relationship graph is curated, not learned. It is auditable and it is small. It does not discover relationships.
- Country-level resolution. The event graph is directed country dyads, not companies. Company-level events arrive through a separate and narrower path.
- Ordering and co-occurrence, never causation. A propagation path is a sequence of observed activations. A shared cause produces the same pattern.
- Market data is research-licensed only. Every price based analysis on this site is research, disclosed as such, and is not redistributed.
- The live ledger is small. No rate is quoted below 30 independent episodes, and correlated calls from a single day count once.