Plain-English methodology. Every claim on this page is either measured and shown with its interval, or marked as not established. Nothing is asserted.
Four things the product does, and their current status. The status is the point: a reader should be able to tell in one line which parts are evidenced and which are not.
| Attention ranking - which record to read first | Withdrawn: no advantage over sorting by event volume |
| Exposure attribution - which holdings an event reaches | Helps a reader engage, but not shown that being correct is why |
| Sentiment score - how adverse this news is for a holding | Withdrawn: the study that supported it measured our own generated text |
| Trading on the ranking - does it make money | No, and five preregistered studies say so |
| Directional calls - will this holding go up or down | Withdrawn, was never distinguishable from a coin flip |
SigmaLens answers two separate questions. Which of today's records deserves an analyst's attention first is answered by the VSigma ranking, which is what the research view orders by and is measured below - ranking was removed from the main product surface after the measurement on this page found no advantage over sorting by event volume. How adverse is this news for this holding is answered by the sentiment score, explained further down. They are different questions and they are measured separately.
How it was measured. A blinded panel of AI evaluators (not human professionals) - five independent AI runs, each instructed to judge from a different professional-investment standpoint - were shown one system's ordering of the day's candidates and asked which three they would look at first. What counts as material was scored by different evaluators who saw no ranking at all, so the target cannot depend on the system being tested. The dates were never used to build the engine.
Every measured comparison is shown above, including the smaller one and the one that did not reach significance. Reporting only the more flattering number would be selective, and this page exists so a reader does not have to take our word for anything.
| Measured on the configuration we ship, 45 blinded AI-evaluator runs (not human professionals) | |
| Block 1, 2025 to 2026 | +0.138, p = 0.276, won 9 of 14 |
| Block 2, 2023 to 2024 | −0.253, p = 0.091, won 5 of 14, and the volume sort won |
| Block 3, crowded days | −0.024, p = 0.847, won 6 of 14 |
| All three combined | −0.046 mean, Fisher p = 0.261 |
| Sign test over every day | 20 of 42, p = 1.00 |
| Superseded, measured on a configuration never shipped | |
| the same three blocks, earlier harness | +0.213 / +0.191 / −0.037, combined p = 0.036 |
The ranking claim is withdrawn. Not reduced, withdrawn. On the engine we ship, ordering by VSigma is not distinguishable from ordering by how many events an entity has.
How the earlier number
happened. The evaluation harness built its exposure graph by calling
DiffusionService() with no entity map; the product builds it with
one. Exposure feeds the ranking through channel relevance, so the harness had
been measuring an engine that was never shipped. Nobody decided this: it is
what happens when the research path and the request path construct the same
object differently and nothing compares them. We found it while auditing an
unrelated coverage change.
How big the correction was. Rebuilding on the shipped construction changed the candidate set on all 42 days: entities the engine correctly declines as unmappable dropped out and newly mapped ones came in. That is why the frozen materiality scores could not be reused and all five gold lenses were re-run for every block, as the preregistration required.
What stops it recurring. The product now hashes its exposure schema, records the hash each result was measured under, and states any mismatch automatically. The superseded blocks stay in the repository with their raw evaluator output so both records can be checked against each other.
What is not affected: the engine still declines to rank records it cannot defensibly score, still shows the evidence behind each reading, and the comparison against the previous ranking engine (+0.046, byte-identical content, order the only difference) stands on its own terms. None of those is a claim that the ordering beats a naive alternative, because that claim is gone.
SigmaLens routes an event to the holdings it reaches and names the channel. The question for a reader who already has a news terminal is whether that improves the read they produce. It has been tested three times on the same twelve events with the same blinded judges, and the answer got more honest each time.
| Run 1: attribution addressed to the entity, not the reader | +0.217, p = 0.42, 6 of 12 |
| Run 2: addressed to the reader's book, weak control | +1.000, p = 0.001, 11 of 12 |
| Run 3: same, with a control that is genuinely wrong | +0.400, p = 0.24, 8 of 12 |
| Our correct attribution vs confidently WRONG attribution | −0.150, p = 0.56, ahead on 3 of 12 |
The run 2 result is withdrawn as a claim about our content. Its control was not a control: three of twelve blocks in the comparison arm were byte-identical to the arm they were meant to test. Rebuilt so the control asserts the opposite of the truth inside the reader's own book, the correct attribution and the wrong one become indistinguishable.
The reason is visible in what the readers wrote. Every one of the twelve given inverted attribution caught it. "It routes crude supply to SPY, GLD and FXI while marking XLE, CVX and USO not reached; that is backwards." They audited the block against their own knowledge, rejected what was wrong, and produced the correct read anyway. For a reader who can audit it, the block works as a prompt to think at the level of the book, and its correctness is not what carries the benefit.
What survives, in these words: a book-addressed attribution block helps a professional reader engage with their own positions. We cannot show that our attribution being correct is what produces that benefit.
One genuinely good finding. Deliberately inverted attribution misled nobody: twelve readers of twelve caught it. That bounds the shipping risk for readers who can audit the claim, and says nothing about a reader who cannot, which is the more dangerous user and one we have not measured.
The ranking numbers above were measured on a configuration that is not the one running. We are saying so because a buyer's diligence would find it, and it should not be the buyer who finds it.
Exposure feeds the ranking through channel relevance. When coverage was extended on 2 September 2026 the ordering moved: on the same 42 evaluation days rank 1 held on 86% of days but the top-three set changed on 52%, and the top three is exactly what evaluators were asked for.
A second, smaller mismatch predates that and we found it while checking the first: the evaluation harness builds its exposure graph without the entity map the product uses, so 18 of 26 entities reached an instrument under the measured configuration against 21 of 26 under the shipped one. Across 14 days that difference is small, rank 1 identical on 14 of 14 and the top three on 13 of 14, and it is disclosed because it is real rather than because it is large.
The product now hashes the exposure schema, compares it with the hash each result was measured under, and states any mismatch automatically, so this cannot go unnoticed again. The fix is to re-run the three blocks against the shipped construction.
Fifteen of the twenty-six entities in the scan universe used to reach no investable instrument at all, so seven of the twelve attribution cases returned "nothing reaches your book" when the truth was "nothing was mapped". On the graph the product actually serves, coverage now reaches an instrument for 21 of 26; two resolve to a real channel carrying no reviewed instrument and say so, and three remain unmapped because no published source establishes them as a material participant in any channel. Mapping them anyway would put a confident path where there is no evidence, which is the one failure this layer exists to avoid.
This page used to say 18 of 26, and five unmapped. Those were the figures for an exposure graph built without the curated entity map, which is not the graph the request path builds. It is the same construction mistake that invalidated the ranking evidence, found a second time during the claims audit of 3 September 2026 and in the one function that exists to report coverage. The coverage helper now defaults to the shipped construction and returns the name of the construction it measured, so a figure of this kind cannot be produced again without saying which engine it describes. The three unmapped entities are GBR, FRA and PAK, and the abstention census confirms them independently: they are declined on all 1,064 days of the panel.
Each candidate is scored on the size of its largest bilateral relationship, corroborated by how much coercive activity it carries and whether it has an evidenced economic channel, with every term compared against the other candidates that same day. Ranking is a comparison, so a record has no rank on its own. Records the engine cannot defensibly score are left out of the order rather than ranked low, and the count of what it declined is shown to the user.
It does not produce alpha, and we tested that. A preregistered study over 1,064 trading days took the top-ranked entities into their mapped instruments and measured excess returns against the market. The result was not distinguishable from zero at any holding period from one to twenty-one days, it did not beat a naive event-volume sort, and it did not survive transaction costs. Four earlier preregistered attempts on the same data also found nothing. (The exact return figures are withheld here, not softened: they are derived from market data whose terms do not permit redistribution. The finding they support, no tradeable edge, checked five times, is stated in full.)
That is not a contradiction of the ranking result. Materiality is not predictability: an event can be genuinely important, correctly ranked first, and already fully in the price. The engine holds no market data and forecasts no direction.
On 1 September we published that the severity score ranks news by materiality, at a rank correlation of about +0.48 against five blinded evaluators. A claims audit two days later found the study could not support it, and the figure is withdrawn rather than restated.
What was wrong. All 120 cases in that study carried a "headline" the system had composed itself, of the form Commodity coverage (Crude Oil): 22 article(s), avg tone -1.5. Separately, the evaluators were shown the source count and the reliability tier, which are two of the three inputs the severity score is built from. So the study asked whether a person reading our summary of our own inputs agrees with our arithmetic over those same inputs. It cannot separate the score from its own description.
What has changed since. Real published titles have been recovered from the source records and 98% of coverage rows now carry one, so the study can be run properly on text the system did not write. It has not been run yet. Until it is, this page makes no claim that the sentiment score ranks by materiality. Everything below describes how the score is computed and what it is not.
The SigmaLens sentiment score (−10 to +10) measures how much news pressure is currently bearing on a holding and which way it points: negative = bearish news pressure, positive = bullish, 0 = quiet. It is built from the geopolitical, commodity-market, and company-specific events we track.
Direction is per holding, not per headline. Adverse geopolitical news is bearish for most assets - but bullish for crisis beneficiaries: oil funds and energy producers when supply is threatened, defense names during conflict, gold as a safe haven. For those holdings the sign flips, and it flips back on de-escalation (a ceasefire is bearish for oil and defense). Every holding's exposure list shows when this rule applies.
It is an informational signal, not a prediction of price or returns, and not a recommendation to buy, sell, or hold. A score near 0 means "quiet in the news we track," not "safe." It reflects news only - it does not model macro drivers like inflation, interest rates, central-bank actions, or earnings.
Each news event gets a strength and a direction, then it is routed to your holding through its exposures:
The event is connected to your holding through its exposures: home country, sector, industry, commodity channels (oil, gas, grain market coverage), the fund's underlying holdings (for ETFs and mutual funds), and the company's own name in global coverage. Your portfolio's net sentiment is the position-weighted average.
We replay the score across a labelled library of major events (2011-2025) and compare it to each asset's benchmark-adjusted move. Two different questions can be asked of that, and for a long time we published the answer to the wrong one.
The wrong question: did a big move follow? That was the original test, and the answer is unimpressive: firing a signal barely selects for a large move at all. But SigmaLens does not claim a big move is coming. It claims a direction: this news is bullish or bearish for this holding. A signal can be exactly right about direction and score nothing on a test that only rewards size.
The right question: was the direction correct? Measured below, against the base rate, because roughly half of all moves are positive and any number near that is noise.
Where does the direction come from, then? Not from that backtest, and we will not pretend otherwise. Direction is a stated rule, not a statistical discovery: an oil-supply shock lifts crude and the companies that sell crude, and squeezes the airlines that buy it. That is market structure a commodities desk would call obvious, and it is why a single market-wide "risk" number is wrong for a portfolio in the first place. What SigmaLens does is apply that rule consistently, per holding, at the speed news arrives.
Because it is a rule rather than a claim, it is auditable: every holding shows the exposures it is scored through and whether each one is treated as a normal asset or a crisis beneficiary. If you disagree with a stance, you can see it and say so. And we are steadily replacing the rule with measurement - each live signal is checked against what the market actually did next, and where there is enough evidence the assumed stance is overridden by the measured one. That is the learning loop below, and the drill-down marks which stances are already measured rather than assumed.
This is the record of a claim SigmaLens used to make and no longer makes: 67.8% directional accuracy on a hand-picked event library. It concerns a scoring approach the product does not ship. It stays here in full because withdrawing a number quietly is the same as never having withdrawn it, but it is folded away because it is history rather than the current record.
The number this page used to lead with was 67.8% directional accuracy against a 54.3% base rate. It came from a library of 83 events: Crimea, Brexit, the 2015 China rout, Soleimani. Those events were chosen because they are memorable, which means they were chosen knowing they moved markets. Measuring how well a model called events selected for being consequential says nothing about whether it calls events.
So we rebuilt the test with nobody choosing. Batches are sampled from the GDELT archive on a schedule fixed before any result was seen, every event runs through the same code that runs in production, and every signal is resolved against the actual benchmark-adjusted five-day move.
| Curated library (what we used to publish) | 67.8% vs 54.3%, n=171 |
| Point in time, 40 weeks | 56.0%, 50 calls from 15 articles |
| Point in time, 104 weeks | 52.5%, 162 calls from 46 articles |
| Point in time, 260 weeks | 49.6%, 401 calls from 116 articles |
| Point in time, 520 weeks | 50.9%, 407 calls from 124 articles |
| Point in time, 104 weeks at six batches a day | 47.6%, 1734 calls from 499 articles |
| Out of sample, 2022 to 2024, never used for research | 48.3%, 2015 calls from 636 articles |
| Edge against a coin flip on the densest sample, 1.7 pp standard error | −2.4 pp, not distinguishable |
| Return per call, five days, net of 10 bp each way | −0.12% (t −0.74), out of sample −0.34% (t −3.28) |
On more than three times the data the edge did not appear. No run is distinguishable from a coin flip. What the money question shows is starker: the calls lose 0.12% each over five days on the densest sample and 0.34% each on a window never used for research, net of costs. The reason is visible in the calls themselves: the engine never called both directions on any holding. Oil and defence names are always long, index funds always short. Its accuracy equals a rule that always says the same thing, because it is a rule that always says the same thing.
A correction to this table, found in our own arithmetic. Until August 2026 the rows above counted every article-and-holding pair as an independent observation. They are not. One wire story naming Russia and Ukraine maps both to the same basket, and GDELT codes that story under several actor pairs, so a single article arrived as up to four identical rows for one holding. That was 47% of the five-year sample and 55% of a two-year re-run: not extra evidence, the same evidence stamped repeatedly. Counting the copies separately inflated the sample and shrank the standard error by roughly 1.8 times, which is enough to turn ordinary noise into a finding. The table now removes the duplicates, and the standard error is resampled by article rather than by row. The conclusion did not change. The precision did, and the ten-year figure changed sign, which is what an error of that kind does.
A third correction, and it was ours to find. Until this audit the base rate in the table above was the realised majority direction of the very sample being scored. That is circular twice over. Nobody knows the majority direction in advance, so nothing can honestly be measured against it; and for a caller that always says the same thing the hits equal the ups, so the base rate equals the accuracy and the edge is zero by arithmetic however good the caller is. On a sample that drifted the other way the same formula invented large negative edges out of the drift. An earlier version of this page reported −5.7 pp as the engine being significantly wrong; that number was measuring the window. The baseline is now a fixed coin flip, the sample's own up-rate is reported separately as context, and the primary metric is the return per call, whose null is zero and needs no baseline at all.
A second correction, found in the same audit. The benchmark was inside its own test universe. A signal raised on the benchmark cannot be benchmark-adjusted, so its abnormal return is 0.000 by construction, and because zero is not greater than zero every bearish call on it was scored correct. Index funds are almost always called bearish, so the contamination was one-sided as well as large: 141 free wins in 1734 rows on the two-year dense run, and between 3.5% and 5.8% of every earlier run. Those rows are now dropped, because scoring the benchmark on its raw return instead would apply a different outcome definition to one holding, which is a second bias rather than a fix. Every figure in the table above is stated after both corrections.
So the directional claim is withdrawn, not softened. What is left is measurement: what a holding has actually moved with, which is checkable, and published here with its own limits.
A third hypothesis, and the one that came closest. Being right and making money are different questions: a signal can be right less than half the time and still pay, if the winners are larger. Traded as a book - long on a bullish call, and variously short, sold or ignored on a bearish one - five years of point-in-time calls returned +0.29% a week above the benchmark. Ten years took that to +0.19%, the t-statistic fell from 1.00 to 0.87, and every variant went negative once a 10 basis point round trip was subtracted. Estimates that shrink as the sample grows are noise decaying, not an edge thinning.
One thing there did survive and is worth stating because it is actionable: the mean move after a bearish call is +0.15%. The name rose. Acting on a bearish call in either direction costs money, so the product does not present one as something to act on.
A second hypothesis, tested and also rejected. If the direction is constant, the intensity might still carry something: news unusually loud for its own entity could precede larger moves whichever way they go. On 37 observations that looked strong, a +4.1% median abnormal move above 1.5 standard deviations against roughly zero otherwise. Extending the same mechanical test to five years took that bucket to 108 observations and the figure collapsed to +0.58%, with the wider above-one-standard-deviation bucket at 0.00% on 166. It was sampling noise, which is what the standard error said at the time. Both candidates have now been tested against five years of point-in-time data and neither survived.
GDELT gives a machine-readable event code, an actor and a conflict score. It does not give the sentence. “Iran seizes tanker in Strait of Hormuz” and “Iran condemns US statement on shipping” produce the same code, the same actor and a near identical conflict score. One is a supply disruption. The other is a press release. So we built a reader that fetches the article's own headline and separates the two, and asked whether the signals backed by an action are more accurate than the signals backed by talk.
The decision rule was written down and committed before the run
returned, in PREREGISTRATION.md, and applied by a script rather than
by a person who had seen the answer.
| Articles read | 499 from 3047 archive batches |
| Backed by an action | 48.0%, 542 calls from 142 articles |
| Backed by talk | 49.3%, 270 calls from 76 articles |
| Difference, action minus talk | −1.3 pp (95% interval −11.3 to +8.6) |
| Two-sided p | 0.80 |
The difference is not merely insignificant, it points the wrong way. It also holds under the earlier, cruder version of the classifier, which we re-ran on purpose because part of that classifier was repaired after we had seen bucket-level returns: −2.1 pp, p 0.68. Reading the article did not help.
What this experiment cannot say. It was powered for a 7 point effect on a single bucket, and the hypothesis it actually tests is a difference between two buckets, whose error is larger and is governed by the smaller of the two. This run could only have resolved a difference of about 13.7 points. So it is honest to say there is no effect of that size, and dishonest to say there is no effect. A smaller one would need roughly 369 articles in the smaller bucket, about four times this sample.
Company filings tell us what a company disclosed. An 8-K item code says results were released; it does not say whether they were good. So we built a parser that reads the earnings release attached to the filing and picks out guidance language - “raised full-year guidance”, “lowered its outlook”, “withdrew prior guidance”. Then we tested whether that predicted anything.
| Filings examined (out of sample) | 227 across 26 companies |
| Filings where the parser called a direction | 58 |
| Directional accuracy, next session, benchmark-adjusted | 58.6% |
| A rule that simply always says “up” | 56.9% |
| Edge, against a 6.6 pp standard error | +1.7 pp |
| Calls that were bullish | 95% (55 of 58) |
The last row explains the rest. Press releases are written to sound positive, so a phrase matcher over them is a bull detector - and in a period when 57% of these names rose, a bull detector scores 57% while knowing nothing. There is a good reason to expect this: “raised guidance” is public within seconds, and prices react to guidance versus expectations, not to the direction of the revision. The worst misses in our test were companies that raised guidance, beat expectations, and fell more than 14%.
So we ship the description and not the direction. SigmaLens will tell you a holding raised guidance, because that is a fact you could not otherwise see on the day it was filed. It will not turn that into a bullish score, because we looked for the edge that would justify doing so and did not find one. The parser was tested on companies whose filings we had never read while writing it; testing it on the four we did read would have measured our ability to fit four documents.
Oil producers gain when supply is threatened, so they carry an inverted stance: adverse news reads bullish for them. That inversion was measured on discrete geopolitical supply shocks, and it stays. What it was also being applied to, and should not have been, was the average tone of general oil-market coverage.
A user asked how a holding's top opportunity could have an average tone of −4.0. Three things were wrong, and none of them were a display bug:
| Commodity channels in the validated backtest | 0 (83 episodes, 11 country entities, none of them a commodity feed) |
| What the channel actually aggregates | oil spills, price crashes and OPEC cuts, averaged into one tone |
| Holdings reached by one such reading | every energy name in the book, at an identical score |
Negative tone in oil coverage is not the same as dearer crude. It is equally the sound of a demand-driven price collapse or a refinery fire, both of which are bad for a producer. The channel had borrowed a result from a test it was never part of, and then applied it to every energy holding at once.
So the channel keeps its volume and loses its direction. It can still tell you oil is being written about loudly, which is true and worth knowing. It no longer turns that into a bullish score. A genuine supply event still reaches an oil producer, via the country-level entities that were tested, and a scenario you ask for still states its own direction because you supplied it. Turning the news channel's direction back on requires a measurement, and there is a flag in the code that will not move without one.
SigmaLens does not just score today's news. Every strong signal it fires is logged, then checked against what the market actually did next (benchmark-adjusted). Those real outcomes retrain the direction model, so the stance for each holding stops being a preset and becomes something measured. This dataset compounds the longer SigmaLens runs.
The events SigmaLens scored strongest, and what the mapped assets actually did (ranked by signal strength, not by outcome - so this is not a cherry-pick):