{"statuses":["SUPPORTED","EXPLORATORY","UNPROVEN","WITHDRAWN"],"counts":{"WITHDRAWN":6,"UNPROVEN":17,"SUPPORTED":2,"EXPLORATORY":4},"claims":[{"id":"ranking_vs_volume","claim":"The VSigma ordering puts more material records at the top than sorting by event volume does.","status":"WITHDRAWN","figure":"-0.05 mean rank points over 42 days, Fisher p = 0.26, won 20 of 42 (sign test p = 1.00)","test":"Three preregistered blocks, 42 days never used to build the engine, 45 blinded AI-evaluator runs (five AI models, each instructed to judge from a different professional-investment standpoint - not human professionals), gold materiality scored by separate evaluators who saw no ordering at all.","source":"research/vsigma/SHIPPED_REPORT.md","withdrawn_figures":["+0.213","+0.191","+0.20","+0.12","p = 0.0137","p = 0.011","p = 0.036"],"why_withdrawn":"The published figures came from an evaluation harness that built its exposure graph without the entity map the product uses. They never described the shipped engine.","safe_to_say":"Measured on 42 preregistered days against a blinded panel of AI evaluators judging from professional investment standpoints (not human professionals), our ordering is indistinguishable from ordering by event volume."},{"id":"directional_accuracy","claim":"The sentiment score predicts whether a holding goes up or down.","status":"WITHDRAWN","figure":"67.8% against a 54.3% base rate on a curated library (n = 171); 47.6% on the densest point-in-time sample (1,734 calls from 499 articles) and 48.3% out of sample","test":"Point-in-time replay in which nobody chooses the events: batches sampled from the archive on a schedule fixed before any result was seen.","source":"templates/methodology.html#withdrawn","withdrawn_figures":["67.8%","13.5 points","54.5%"],"why_withdrawn":"The 67.8% came from a hand-picked event library. Measured mechanically it is not distinguishable from a coin, and trading the calls loses money net of costs.","safe_to_say":"We used to publish a directional accuracy figure. It did not survive a point-in-time test and we withdrew it. We make no directional claim."},{"id":"ranking_alpha","claim":"Trading on the ranking produces excess return.","status":"UNPROVEN","figure":"-0.0010 excess return per five-day trade, 1,064 days, five preregistered studies, none positive","test":"Preregistered event studies against benchmark-adjusted returns on research-licensed price data.","source":"research/alpha/PREREGISTRATION-VSIGMA-ALPHA.md","safe_to_say":"We have looked for a return effect five times and found none. We do not claim one. Materiality is not predictability."},{"id":"predictive_incremental","claim":"Some layer of the SigmaLens representation (event volume, a sentiment proxy, event mechanism, or exposure) adds out-of-sample return information beyond ordinary market/sector context.","status":"UNPROVEN","figure":"all four nested layers fail at h=5, Bonferroni alpha 0.0125: volume p=0.879, valence proxy p=0.150, mechanism p=0.950, exposure p=0.483; holdout and 14 leave-one-out break tests show the same pattern","test":"Preregistered nested ridge-regression walk-forward, 994 signal days x 14 instruments, expanding window with a 21+h day embargo, stationary block bootstrap, holdout read once after the design was frozen.","source":"research/predictive/PREREGISTRATION-INCREMENTAL.md","safe_to_say":"We tested whether event volume, a sentiment stand-in, event mechanism, or exposure each add return information beyond plain market context. None does, at any of the four horizons tested, in the full sample, in the untouched final holdout, or in any of 14 sector/year robustness cuts. This is the sixth preregistered attempt to link this panel to returns; all six have failed."},{"id":"event_driven_alpha","claim":"A specific SEC-filed corporate event (an insider's open-market purchase, an activist Schedule 13D stake, or an 8-K item with a declared bullish/bearish prior) predicts tradeable excess return in the affected security.","status":"UNPROVEN","figure":"zero of four preregistered tests clear Bonferroni at h=5: insider purchases p=0.615 (wrong sign), purchases-vs-sales p=0.886, Schedule 13D p=0.041 (excludes zero but the WRONG sign), declared-direction 8-K p=0.703","test":"Preregistered event-level panel, 18,761 rows, 30 DJIA constituents, 2021-10-01 to 2026-06-30, SEC's own bulk Form 3/4/5 dataset plus the submissions API, holdout read once from 2025-06-12.","source":"research/events/PREREGISTRATION-EVENTS.md","safe_to_say":"We tested the two most-replicated informed-trading signals in the finance literature -- insider open-market purchases and activist 13D stakes -- plus the product's own declared 8-K direction prior, against SEC filing data for the 30 largest US companies. None predicted tradeable return. Where one result reached significance before correction, it pointed the wrong way. This does not test smaller, less-covered companies, where the same literature expects a larger effect and where the question remains open."},{"id":"company_incremental","claim":"Some layer of SigmaLens's company-level event representation (event volume/attention, the shipped valence/direction, event mechanism, or exposure/materiality) adds out-of-sample return information beyond ordinary market/sector context, on real SEC filings across a broad universe.","status":"UNPROVEN","figure":"all four nested layers fail at h=5, Bonferroni alpha 0.0125, every one wrong-signed: event volume p=0.103, valence p=0.409, mechanism p=0.916, exposure p=0.274; OOS R^2 and IC decline monotonically as layers are added (M0 -0.0022 -> M4 -0.0049); holdout spans zero for all five models","test":"Preregistered nested ridge-regression walk-forward, 63,087 (ticker, day) rows, 503 S&P 500 constituents, 1,250 unique event-days, 2021-10-01 to 2026-07-15, real Form 4/8-K/13D-13G filings via analysis/corporate.py's shipped classification (not a research proxy), expanding window with a 21+h day embargo, stationary block bootstrap, holdout read once after the design and split were frozen.","source":"research/company_incremental/PREREGISTRATION-COMPANY.md","safe_to_say":"We tested whether event volume, the product's own shipped directional read, event type/mechanism, or trade-size materiality each add return information beyond plain market context, on real SEC filings across 503 large-cap companies -- about 17x the cross-section and roughly 30x the insider-purchase sample size of our first company-level test. None does. Every layer's effect is not just indistinguishable from zero but points the wrong way, and forecast accuracy gets worse, not better, as more of the representation is added. This is the seventeenth preregistered attempt in this repository to link an event panel to returns; all seventeen have failed. The one untested dimension left is smaller, less-covered companies."},{"id":"attribution_helps","claim":"Telling a reader which of their holdings an event reaches helps them more than the news alone.","status":"UNPROVEN","figure":"correct attribution vs confidently wrong attribution -0.150, p = 0.56, ahead on 3 of 12","test":"Three runs of blinded readers. Run 3 added a control in which the attribution was inverted inside the reader's own book while keeping the same form and vocabulary.","source":"research/attribution/ATTRIBUTION_REPORT_RUN3.md","safe_to_say":"Book-addressed attribution beats no attribution, but against a control that was confidently wrong we could not show the correctness is what helps: readers audited and corrected it either way. Format and content are still confounded."},{"id":"sentiment_ranks_materiality","claim":"The severity score ranks news by materiality.","status":"WITHDRAWN","figure":"rho +0.483 against blinded materiality, 120 events","test":"120 events scored by five blinded AI evaluators (not human professionals) shown the headline, category, entities, date, source count and reliability tier.","source":"research/sentiment/SENTIMENT_REPORT.md","withdrawn_figures":["+0.483","rho +0.48"],"why_withdrawn":"Audited 3 September 2026. All 120 of those 'headlines' were prose SigmaLens generated itself ('Commodity coverage (Crude Oil): 22 article(s), avg tone -1.5'), and the evaluators were shown the source count and reliability tier that the severity score is computed from. The study could not separate our score from our own summary of its own inputs. Real published titles have been available since the GKG column 26 fix; the study has not been re-run on them.","safe_to_say":null},{"id":"confidence_percentage","claim":"The confidence percentage tells a reader how much to trust a reading.","status":"WITHDRAWN","figure":"100% confidence on any single SEC or OFAC filing, regardless of how many outlets reported it","test":"Every scored entity-day whose underlying evidence was a single regulatory filing: the source-tier-and-count formula assigns a filing a corroboration count of 16 by construction, which alone saturates the log curve to 100%.","source":"analysis/evidence_quality.py","withdrawn_figures":["100% confidence"],"why_withdrawn":"It does not discriminate, and on a single SEC filing it read 100% because a filing is assigned a corroboration count of 16 by construction. The product now states the evidence itself: what kind of source, how many, whether anybody counted them, how old, and the route to the holding.","safe_to_say":"The old confidence percentage was the source-tier and corroboration count shown alone; on a single SEC or OFAC filing that construction put it at 100%, so we withdrew it. What replaced it, in analysis/confidence.py, is a differently built score: the old count is now one of six checkable factors (the others being signal/macro/sector agreement, track record, and contradicting evidence), under a 90-point ceiling because nothing here reaches certainty, with absence of a factor never raising the score."},{"id":"abstention_rate","claim":"The system declines to score what it cannot defensibly score, and says why.","status":"SUPPORTED","figure":"declines 43.1% of entity-days (11919 of 27664 over 1064 days); 3192 of those are 3 entities with no mapped exposure at all, declined every day","test":"A census, not an estimate: every day in the frozen panel, the shipped engine ranking the standard 26-entity scan universe, counting what it scored and what it declined together with the reason it gave.","source":"research/abstention/COVERAGE.json","safe_to_say":"Across 1064 days of the frozen panel the engine declined 43% of the entity-days it was asked about, and every decline carries the reason it gave. That figure is two things: FRA, GBR, PAK have no mapped exposure at all and are declined every single day, which is a coverage gap and not restraint; on the entities we do cover the rate is 36%. It is a count of our own behaviour either way, not a performance claim, and we have not shown that the declined records deserved to be declined.","detail":{"days":1064,"entity_days":27664,"declined":11919,"abstention_rate":0.4308,"never_scored":["FRA","GBR","PAK"],"structural_declines":3192,"evidence_declines":8727,"evidence_abstention_rate":0.3566,"top_reason":"only N events, below the minimum of N","top_reason_share":0.753,"reasons":[{"reason":"only N events, below the minimum of N","count":8975,"share_of_declines":0.753,"entities":[["LBY",1017],["YEM",955],["VEN",831],["IRQ",767],["PRK",698],["TWN",645]]},{"reason":"exposure is UNMAPPED and no evidenced economic channel for this country, so no defensible relevance exists","count":3192,"share_of_declines":0.2678,"entities":[["GBR",1064],["PAK",1064],["FRA",1064]]},{"reason":"no activity in the window","count":1113,"share_of_declines":0.0934,"entities":[["LBY",463],["YEM",228],["VEN",126],["PRK",98],["IRQ",66],["LBN",50]]}]}},{"id":"confidence_discriminates","claim":"Records the engine is more confident about are the ones blinded professionals find more material.","status":"EXPLORATORY","figure":"salience rho +0.67 against blinded materiality; the event-volume baseline gets +0.66 on the same rows (n = 420 over 42 days)","test":"The engine's band and salience against the blinded gold materiality already collected on the 42 shipped-configuration evaluation days, with the null permuted within each day. Event volume is scored the same way on the same rows.","source":"research/abstention/CONFIDENCE_VS_GOLD.json","safe_to_say":"Exploratory, on evidence collected for another experiment. The engine's salience is rank-correlated with blinded professional materiality at rho +0.67 across 420 candidates on 42 days. Counting events on the same rows gets rho +0.66, which is the same thing to two decimal places: the score is not noise, and it is not better than counting. The stated confidence LABEL is a separate matter and it does not discriminate: it reads HIGH on almost every scored record and correlates at only rho +0.12. What carries the uncertainty in this system is the decision to decline, not the label.","detail":{"band":{"rho":0.5827,"p":0.0,"n":420,"days":42},"salience":{"rho":0.6669,"p":0.0,"n":420,"days":42},"confidence_label":{"rho":0.1238,"p":0.002,"n":419,"days":42},"volume_baseline":{"rho":0.66,"p":0.0,"n":420,"days":42},"gold_by_band":{"HIGH":{"n":123,"mean_gold":6.0022},"CRITICAL":{"n":35,"mean_gold":7.24},"MEDIUM":{"n":114,"mean_gold":4.9158},"LOW":{"n":148,"mean_gold":4.2284}},"gold_by_confidence":{"MEDIUM":{"n":13,"mean_gold":4.0},"HIGH":{"n":406,"mean_gold":5.2225}}}},{"id":"exposure_coverage","claim":"The exposure graph says which holdings an event reaches, and which entities reach nothing.","status":"SUPPORTED","figure":"21 of 26 scan entities reach an instrument; 2 macro only; 3 UNMAPPED","test":"A count over the curated schema, with a configuration fingerprint compared at serve time so the served graph and the measured graph cannot drift apart silently.","source":"analysis/exposure_graph.py","safe_to_say":"21 of the 26 entities we scan reach a tradable instrument through a documented channel with a citation and an effective date. 2 reach only a macro channel and 3 reach nothing at all and are reported as unmapped, rather than being given a relationship we cannot evidence. The 3 unmapped ones are never scored on any day.","detail":{"fingerprint":"99156753b605","construction":"shipped","entities":26,"reach_an_instrument":["CHN","IND","IRN","IRQ","ISR","JPN","KOR","LBN","LBY","NGA","PRK","PSE","RUS","SAU","SYR","TUR","TWN","UKR","USA","VEN","YEM"],"macro_only":["DEU","EGY"],"unmapped":["FRA","GBR","PAK"]}},{"id":"live_record","claim":"The signals it has issued since launch have been right more often than the base rate.","status":"UNPROVEN","figure":"51.5% across 306 resolved episodes since launch","test":"Every signal is written to a ledger and resolved against the benchmark-adjusted move. No rate is quoted below 30 independent episodes.","source":"storage/sqlite_db.py","safe_to_say":"Since launch the live ledger reports 51.5% across 306 independent episodes. It is a small sample and it is reported separately from any backtest."},{"id":"secondorder_exposure_alpha","claim":"SigmaLens identifies second-order exposures -- a country shock propagating through a curated exposure channel to a dependent security -- that the market prices more slowly than the directly affected security itself.","status":"UNPROVEN","figure":"the Taiwan/China/Korea channel produced zero qualifying adverse episodes in 1,441 days (insufficient data, not tested); the oil channel's confirmatory p-values (n=15) are a bootstrap-implementation artifact, not evidence -- the one statistic computed without that artifact (exposure-weight x direct-move interaction) is negative and non-significant, coef=-0.138, p=0.11","test":"Preregistered GDELT 1.0 daily event panel, 2022-07-01 to 2026-06-30, two curated exposure channels reused verbatim from analysis/exposure.py, causal rolling z-score episode construction, episode-level stationary block bootstrap, holdout read once.","source":"research/secondorder/PREREGISTRATION-SECONDORDER.md","safe_to_say":"We tested whether a second-order exposure channel prices in more slowly than the directly affected security. One of two channels (Taiwan/China/Korea) never produced a usable adverse episode -- the country-level event mix there is structurally cooperative-coded, not a data-collection failure. The other (oil) reached only 15 qualifying episodes, below the block bootstrap's reliable-resampling floor of 20, so its apparent significance is a coding artifact we did not act on. The one cleanly computed statistic argues against the mechanism, not for it."},{"id":"propagation_alpha","claim":"After a Taiwan/China conflict-flagged information shock, SigmaLens's curated chip-supply-chain dependency weight identifies which secondary securities underreact and then reprice, beating weakly-exposed peers over 1-10 trading days.","status":"UNPROVEN","figure":"0 of 3 preregistered confirmatory hypotheses passed on 44 real episodes (2022-12-12 to 2026-06-10, 33 in train+dev): H1 spread p=0.631, H2 propagation-speed p=0.265, both indistinguishable from zero at every horizon (h=1/3/5/10/20) and every entry delay (T+0/1/2). H3 (weight x direct_move interaction, the one estimator in this repository with no small-N bootstrap defect) is significant, coef=-0.931, p=0.0, n=660 rows/33 episodes -- but the sign is backwards: higher-dependency names co-moved LESS with the direct entity's own return, not more.","test":"Preregistered event-level GDELT construction (QuadClass=4 Material Conflict rows, bilateral actors, NumMentions>=10, rolling z-score on daily qualifying-row count, not a raw Goldstein sum), reusing analysis/exposure.py::SUPPLY_CHAIN weights verbatim. Chronological 50/25/25 split, holdout read once, Bonferroni alpha=0.0167 across 3 hypotheses.","source":"research/propagation/PREREGISTRATION-PROPAGATION.md","safe_to_say":"We tested whether SigmaLens's curated Taiwan/semiconductor dependency graph predicts a delayed, same-direction repricing after a real geopolitical conflict shock. It does not: the spread is flat at every horizon we checked. The one robust, artifact-free statistic in the study found the dependency weight carries real information, but in the opposite direction from the hypothesis -- more heavily-exposed names moved less like the direct entity, not more. This is not a data or construction failure like the prior second-order study; the event and exposure construction held up under sanity checks. It is evidence against the propagation mechanism itself."},{"id":"earnings_surprise_alpha","claim":"A company's standardized earnings surprise (SUE: reported EPS vs. the pre-report analyst consensus, scaled by the firm's own trailing surprise dispersion) predicts its subsequent sector-adjusted return over the following trading days -- post-earnings-announcement drift.","status":"UNPROVEN","figure":"0 of 3 preregistered confirmatory hypotheses passed on 2,367 (ticker, quarter) observations, 67 tickers, 2017-2026 (1,780 in train+dev): H1 (SUE vs. h=3 sector-adjusted return, Spearman) rho=-0.002, p=0.926; H2 (does SUE beat momentum+vol+direction, clustered OLS) SUE coefficient p=0.504; H3 (top-vs-bottom SUE tercile spread) spread=-0.20bp, p=0.491. Holdout (587 observations, read once): H1 rho=+0.020 p=0.600, H3 spread=+0.24bp p=0.499 -- signs flip between train+dev and holdout with neither remotely significant, consistent with pure noise around zero.","test":"Preregistered: yfinance EPS-estimate-vs-actual history, standardized by the firm's own trailing 8-quarter dollar-surprise dispersion (min 4 quarters, min $0.01 denominator -- see PHASE0-AUDIT.md and the mid-study MIN_SUE_DENOMINATOR fix in universe.py). 67-ticker, 9-sector universe, every ticker verified for EPS-fill-rate in Phase 0. Chronological 50/25/25 split, holdout read once, Bonferroni alpha=0.0167 across 3 hypotheses.","source":"research/expectation/PREREGISTRATION-EXPECTATION.md","safe_to_say":"We tested whether SigmaLens's data can detect post-earnings-announcement drift, one of the best-established anomalies in finance, using a free, retail-grade earnings-consensus source across 67 liquid, well-covered companies over 10 years. It cannot: every measure is indistinguishable from zero, in-sample and out-of-sample, with no consistent sign. This does not contradict the academic literature -- it may reflect this universe's heavy analyst coverage and liquidity (a known moderator: drift concentrates in less-covered names), the data source's retail-grade point-in-time quality, or the specific horizon tested -- none of which is evidence a differently constructed test would find something, and none was chased further per this project's standing rule against re-cutting a null."},{"id":"analyst_revision_alpha","claim":"A sell-side analyst's continuous price-target revision (ln of new vs. own prior stated target) predicts a security's subsequent sector-adjusted return over the following trading days.","status":"UNPROVEN","figure":"0 of 3 preregistered confirmatory hypotheses passed on 19,025 analyst-action observations, 66-67 tickers, 2016-2026 (14,268 in train+dev): H1 (REVISION vs. h=3 sector-adjusted return, Spearman) rho=+0.004, p=0.876; H2 (does REVISION beat momentum+vol+direction, clustered OLS) REVISION coefficient p=0.241; H3 (top-vs-bottom REVISION tercile spread) spread=+0.13bp, p=0.615. Holdout (4,757/3,013 observations, read once): H1 rho=-0.088 p=0.003, H3 spread=-0.77bp p=0.006 -- BOTH significant, but opposite in sign from train+dev, and the preregistered stop rule (0 of 3 on train+dev) had already fired before holdout was opened. This is disclosed, not chased: treating the significant holdout numbers as the finding would be exactly the post-hoc-holdout-reversal trap the preregistration exists to prevent.","test":"Preregistered: yfinance's upgrades_downgrades (real dated sell-side rating/price-target-action history, 2012+ for large-caps), restricted to actions with both a prior and a current price target present. Reused the earnings study's 67-ticker, 9-sector universe. Chronological 50/25/25 split, holdout read once, Bonferroni alpha=0.0167 across 3 hypotheses.","source":"research/revision/PREREGISTRATION-REVISION.md","safe_to_say":"We tested whether the size of a sell-side analyst's own price-target revision -- not just whether they upgraded or downgraded, but by how much -- predicts a stock's next few days of sector-adjusted return, using a genuinely new, previously-unused, point-in-time data source (this had never been wired into any SigmaLens pipeline before). It does not: train+dev shows no distinguishable-from-zero effect on any of three tests. Holdout alone shows a significant NEGATIVE relationship, which we are explicitly not reporting as a discovery -- the preregistered decision to stop had already been made before holdout was read, precisely so a result like this can't be quietly promoted after the fact. A follow-up diagnostic audit (research/revision/CAUSALITY-AUDIT.md, not a return study) found the mechanism: REVISION is very strongly reactive to the prior price move it follows (Spearman rho up to +0.64 at 60D; a 10% prior abnormal move predicts roughly a 7.6-8.2% price-target revision in the same direction), and REVISION's own coefficient loses significance once the prior move is controlled for -- consistent with revision being a lagging confirmation of a move already in the price, not an independent information source. No new preregistered hypothesis was opened: the holdout's reversal-like pattern concentrates in exactly the subgroups (extreme prior moves, clustered analyst actions) that this same audit had to inspect to find it, so freezing thresholds on that basis would not be a clean test."},{"id":"multi_source_convergence_alpha","claim":"A security touched by multiple, distinct SEC-disclosure channels (an 8-K, a Form 4 insider transaction, a 13D/G ownership filing) within the same week -- 'convergence' -- shows a larger subsequent price reaction than a security touched by only one channel.","status":"UNPROVEN","figure":"0 of 3 preregistered confirmatory hypotheses passed on 14,068 DJIA-30 filing-event observations, 30 tickers, 2021-2025 train+dev (18,761 total incl. holdout). H1 (convergence_count vs |h=3 sector-adjusted return|, Spearman) rho=-0.001, p=0.94; H2 (does convergence_count beat momentum+vol+severity, clustered OLS) coefficient p=0.523; H3 (HIGH>=2-channel vs LOW=1-channel spread) spread=-0.01bp, p=0.857. The decisive baseline -- raw event VOLUME in the same window, regardless of channel diversity -- was ALSO null (rho=-0.029, p=0.295), so this is not a repackaging of the 1st study's closed event-volume hypothesis working under a new name; it is independently null. Mean |return| is flat (1.5-1.7%) across every one of the 7 possible channel combinations. Holdout (4,693 observations, read once): H3 spread=-0.36bp p=0.008, significant but opposite in sign from train+dev -- disclosed, not chased, per the same preregistered stop rule as the analyst-revision study.","test":"Preregistered: reused the 7th study's own DJIA-30 filing panel (research/events/PANEL.json) verbatim, deriving one new feature never before compared to a return -- the count of distinct disclosure channels (8-K / insider Form 4 / 13D-G ownership) active in a trailing 5-calendar-day window. Chronological 50/25/25 split reused from the source panel, holdout read once, Bonferroni alpha=0.0167 across 3 hypotheses.","source":"research/convergence/PREREGISTRATION-CONVERGENCE.md","safe_to_say":"We tested whether a security being touched by multiple INDEPENDENT SEC disclosure channels at once -- not just one filing, but several different kinds within the same week -- predicts a bigger subsequent move, using a feature that had never been examined before this test. It does not: every one of three preregistered tests came back null, and critically, so did the obvious alternative explanation (it's just more filings of any kind, i.e. volume) -- so this isn't the 1st study's closed event-volume finding wearing a new name, it is a second, independent null. Holdout showed a significant reversal, which we are explicitly not reporting as a discovery for the same reason the analyst-revision study's holdout wasn't: the preregistered decision to stop had already been made before holdout was read."},{"id":"govcon_contract_modification_alpha","claim":"A large federal contract MODIFICATION (follow-on funding on an existing award, as opposed to a new award) on a publicly traded defense/govcon contractor predicts a same-direction move in that stock's price over the following trading days, because modifications -- unlike initial awards -- are not routinely re-announced and so may be underpriced.","status":"UNPROVEN","figure":"1 of 3 preregistered confirmatory hypotheses passed on 9,081 modification-only observations, 25 govcon tickers, 2020-2024 train+dev (20,290 total transactions incl. initial awards and holdout). H1 (signed transaction amount vs. h=3 sector-adjusted return, Spearman) rho=+0.009, p=0.529; H2 (does amount beat momentum+vol, clustered OLS) coefficient p=0.286; H3 (top-vs-bottom signed-tercile spread) spread=+18bp, p=0.014 -- clears Bonferroni (0.0167) narrowly. H3 did NOT replicate in holdout (n=2,404, spread=+14bp, p=0.221). Economic significance on the train+dev spread itself: hit rate 49.8% (chance), and the mean spread turns NEGATIVE net of just 10bp round-trip costs (-4bp). A preregistered diagnostic found modifications get essentially the SAME immediate market reaction as initial awards (mean |entry-day return| 2.16% vs. 2.32%) -- the mechanism's own premise (modifications are less-covered, so underpriced) has no support in this data.","test":"Preregistered after a manual review of the 20 largest modification transactions found the largest, most material ones cluster on a handful of the most heavily defense-press-covered programs in the country (F-35, KC-46, Virginia-class submarines) -- disclosed as a reason to expect a null before the confirmatory run, not discovered after. New 25-ticker govcon universe (SIC classification + 5 named additions, market cap floor), new USAspending.gov transaction-level connector (real per-action dates and per-action incremental amounts, not the cumulative award total). Chronological 50/25/25 split, holdout read once, Bonferroni alpha=0.0167 across 3 hypotheses.","source":"research/govcon/PREREGISTRATION-GOVCON.md","safe_to_say":"We built a brand-new connector to federal contract award data (USAspending.gov) and a new universe of 25 publicly traded defense contractors -- a genuinely new information source and the first universe in this project's history where a single contract can be material to the stock. One of three preregistered tests cleared significance on development data, but it did not replicate in the untouched holdout, the underlying spread would lose money after realistic trading costs even before holdout, and a direct check found the mechanism's own premise unsupported: modifications do not get a smaller immediate market reaction than brand-new awards, which is the whole reason we expected them to be underpriced. No signal."},{"id":"duration_classifier_decay_alpha","claim":"SigmaLens's current production event-duration classification (analysis/horizon.py: immediate/near/structural) contains information about the TIME PROFILE of a security's subsequent abnormal-return decay -- i.e. structural events' price impact persists longer than immediate events'.","status":"UNPROVEN","figure":"0 of 3 preregistered confirmatory hypotheses passed on 839 usable train_dev observations (900 total, 28 tickers), DJIA-30, 2021-2024 8-K events. H1 (primary, single pre-specified omnibus cluster-permutation test of whether the magnitude-controlled decay curve across 7 horizons differs by band) stat=14.07, p=0.147 -- not significant even at conventional 0.05, let alone the required Bonferroni 0.0167. H2 (structural-vs-immediate spread at a 20-day half-life checkpoint) spread=+1.51, p=0.037 -- misses Bonferroni. H3 (band dummies incremental over momentum/vol/severity, clustered OLS) is_structural p=0.085, is_near p=0.104; R2 moved from 0.0096 to 0.0154. A raw (non-magnitude-controlled) version of H1 was also null and less significant (p=0.319), ruling out 'the classifier just tracks event size' as the reason for the null. Holdout (600 obs, read once): H1 stat=41.73 p=0.061, H2 spread=+4.06 p=0.148 -- bigger point estimates in the theoretically expected direction, neither significant. Since H1 already failed on train_dev, the duration-aware trading test was never run, in either direction.","test":"Preregistered after auditing the exact production classifier (analysis/horizon.py) -- a hand-authored lookup table whose own docstring states 'nothing here has been measured yet' -- and confirming it reproduces exactly, deterministically, with no future-information risk. Tested on the one slice with sufficient cached historical data (8-K item-code path, reusing the already-closed 7th study's panel verbatim); the GDELT-sourced country/company slices of the same classifier are reported as genuinely UNTESTED, not passed or failed, for lack of a sufficiently large cached dataset with real theme labels. Chronological 60/40 split (mission-specified, not house default), holdout read once, Bonferroni alpha=0.0167 across 3 hypotheses.","source":"research/persistence/PREREGISTRATION.md","safe_to_say":"We tested whether SigmaLens's existing short-term/medium-term/structural event labels -- which the code itself already documented as unmeasured reasoned priors, not something learned from data -- actually predict how fast an event's market impact fades. On the slice of the classifier we could test with real historical data, they do not, at the bar we set in advance: the single pre-specified test of curve shape came back not significant, and the closest secondary result (p=0.037) missed the multiple-testing-corrected threshold. A manual review found no labeling defects -- the mundane explanation is that ordinary company-specific noise over 20-120 trading days is large enough to swamp whatever real effect this classification implies. This does not touch the GDELT-sourced part of the same classifier, which remains untested for lack of data, not because it passed."},{"id":"event_conditional_relative_value_alpha","claim":"Conditional on the same shock episode, securities SigmaLens's exposure/stance tables classify as MORE POSITIVELY exposed outperform securities it classifies as MORE NEGATIVELY exposed over the subsequent window -- a cross-sectional relative-value spread, not a directional prediction about any single security.","status":"UNPROVEN","figure":"0 of 3 preregistered confirmatory hypotheses passed on the sector_board channel (China/Russia/Taiwan, 128 episodes, 77 train_dev / 51 holdout). H1 (primary, pooled long-short spread at h=5, cluster-bootstrapped) stat=+0.061%, p=0.806 on train_dev -- an order of magnitude below anything economically meaningful, let alone the required Bonferroni 0.00833 (corrected across 6 hypotheses spanning two channels). H2 (severity_z coefficient, controlling for event volume/momentum/vol/entity) p=0.127. H3 (differential-scaling: |severity| vs spread) stat=-0.117, p=0.252 -- wrong sign. A 2,000-shuffle placebo found the real long/short assignment statistically indistinguishable from a random one (p_placebo=0.813). Holdout's H1 point estimate (+0.79%, p=0.045) is an order of magnitude larger and nominally significant but still misses Bonferroni, and per the preregistered stop rule does not count since train_dev already failed. A second channel (oil_producer_consumer, 38 episodes, reusing Study 8's already-cached data) could not be tested to the same standard: its single-pseudo-entity design produces only 8 clusters against a 60/40 split, below the MIN_CLUSTERS=15 floor on both sides -- reported as insufficient sample, not rescued by re-splitting after seeing the shortfall.","test":"Preregistered after a Phase 0 audit of the exposure/stance machinery (analysis/exposure.py, analysis/exposure_graph.py, analysis/scenarios.py::SECTOR_BOARD) found a genuine two-sided (benefit-leg AND hurt-leg) construction for only 5 of 11 curated entities; scoped to the 3 (China, Russia, Taiwan) where neither leg depends on an operationally ambiguous country code. Legs built entirely from already-shipped sector/stance tables, no new exposure relationship invented. A second channel (11-country oil producer-vs-consumer split, reusing Study 8's OIL_UNIVERSE and cached GDELT data) was added mid-study, before either channel's returns were read, once the primary channel's fresh per-country GDELT fetch proved network-bound in this environment. Chronological 60/40 split per channel, holdout read once, Bonferroni alpha=0.05/6 across both channels' H1-H3.","source":"research/relative_value/PREREGISTRATION.md","safe_to_say":"We tested whether SigmaLens's exposure reasoning can tell that the SAME piece of news means different things for different holdings, in a way that shows up as a tradeable long-short spread rather than just a prediction about one stock. On the channel we could test properly (China, Russia, Taiwan), it does not: the primary test came back essentially at zero, a placebo showed the real assignment was no different from a random one, and the one entity (Russia) whose long and short sides are both built from entity-specific reasoning (rather than a generic safe-haven proxy) came out with the wrong sign on average. A manual review found no construction defect -- the likely mundane explanation is that a whole sector's ordinary week-to-week volatility (semiconductors, for the Taiwan leg) swamps whatever differential signal a single news day carries. A second, oil-based channel could not be tested to our own standard at all, for a sample-size reason disclosed rather than patched around."},{"id":"filing_dissimilarity_nowcast_alpha","claim":"The year-over-year textual dissimilarity of a company's own 10-K Item 1A Risk Factors section nowcasts its forward revenue-growth surprise relative to its own trailing trend, incremental to trend continuation alone, and this expectation gap or the dissimilarity itself corresponds to a subsequent abnormal stock return.","status":"UNPROVEN","figure":"Nowcast validation gate failed outright: adding dissimilarity to a trend-only revenue-growth forecast produced a WORSE holdout MAE (1.113 vs. 1.070, 95% CI on the difference entirely negative, [-0.052, -0.036], p=0.000). Of 4 Bonferroni-corrected confirmatory tests (alpha=0.0125) on 968 (company, fiscal-year) observations across 99 companies (91 dev / 98 holdout clusters): H1 (dissimilarity vs. raw forward growth) p=0.914 dev / 0.806 holdout; H2 (vs. trend-relative surprise) p=0.696 dev / 0.781 holdout, sector/tier-controlled coefficient p=0.613; H4 (vs. +252-trading-day abnormal return) p=0.993 dev / 0.647 holdout. Two placebos (shuffle across companies within year; across years within company) both placed the real statistic inside the null band. Leave-one-company/year/sector-out moved the pooled statistic by at most 0.03 in any direction -- a broad, not masked, null. H3 (does a company's REALIZED surprise, known only in hindsight, correspond to a subsequent abnormal return -- a mechanism-validity check, not a test of this signal) DID clear Bonferroni and replicate cleanly: dev stat=0.198 p=0.000, holdout stat=0.211 p=0.001.","test":"Preregistered after an 18-candidate Phase 0/1 audit of genuinely free, point-in-time-auditable information sources this repository had not already tested (see research/money_making/CANDIDATE_UNIVERSE.md), selected over a govcon-contract-backlog alternative that failed on sample size before any data was fetched (research/money_making/SELECTION.md). 99-company universe drawn by a fixed-seed stratified rule (11 GICS sectors x 3 cap tiers) before any 10-K was fetched. Trend-relative expectation used in place of analyst consensus, which is licensed data this product does not have. Chronological split (dev: fiscal year <= 2021, holdout: >= 2022), read once. Two real extraction/labelling bugs were found and fixed on live data before any dissimilarity score existed: a whitespace-collapsed heading matcher (an inline HTML tag splitting a word like 'Risk Factors' mid-string was silently missed by a plain regex) and a fiscal-year collision fix for 52/53-week filers and a transition-period 10-K, both caught by sanity.py's own reportDate-spacing check before run_study.py was invoked.","source":"research/money_making/PREREGISTRATION.md","safe_to_say":"We tested whether the amount a company rewrites its own annual risk-factor disclosure from one year to the next says anything about its future revenue growth or its stock. It does not: the forecast-validation gate failed before any return was examined (the signal-augmented forecast was measurably worse than simple trend continuation), and all three tests of the signal itself replicated a clean, broad null on a well-powered holdout (99 companies, 968 observations -- the best-powered study in this research programme). A manual case review found no construction defect: the same company (Owens Corning) produced the two highest dissimilarity scores in the entire panel five years apart, with opposite subsequent outcomes (+41% and -40%), a working illustration of why a measure of HOW MUCH language changed cannot also say WHICH WAY that change points. One separate, genuine finding survived: a company's realized revenue surprise relative to its own trailing trend does correspond to a subsequent stock move, cleanly and on both data splits -- useful context for a future study, but it does not rescue this one, since it says nothing about SigmaLens's own ability to see the surprise coming."},{"id":"fresh_eyes_ranking_value","claim":"Showing SigmaLens's full scored output (unusualness, diffusion, exposure, historical context, provenance, confidence) helps an investment professional pick more materially useful items than a plain volume ranking of the same underlying event data.","status":"WITHDRAWN","figure":"C (full product output) - B (volume ranking) = -0.383 mean blind materiality points on a 0-3 scale, 95% CI [-0.566, -0.175], excludes zero; 5 of 5 independent evaluator lenses agree on direction; raw unranked evidence (1.696) beat both the volume ranking (1.649) and the scored product output (1.266)","test":"Preregistered blinded fresh-eyes evaluation, 30 independent evaluator agents across 5 lenses and 4 phases, 32 cases and 80 ranking candidates on dates never used to build or tune the engine. Blind materiality judged before any system output was revealed to any evaluator; aggregated once, after all 30 outputs existed. Adequately powered for the preregistered 0.40-point effect (minimum detectable effect 0.286 at 80% power).","source":"research/fresh_evaluation/FINAL_REPORT.md","withdrawn_figures":["+0.40","STRONG PRODUCT VALUE","PROMISING"],"why_withdrawn":"The composite score elevated reporting-frequency churn (wide-partner diplomatic hubs like DEU and GBR whose counterparty mix rotates for routine reasons) over genuinely material records. The same evaluation independently caught, in 8 of 32 cases across all 5 lenses, the ENTITY_MAP bucket defect that labelled Lebanon and Syria DIRECT exposure to crude and defence tickers; that specific defect has since been fixed (see analysis/exposure_graph.py's ASSOCIATED class and GROUP_ENTITIES gate). The propagation hop-count and no-outcome-data analogue-panel defects the same evaluation surfaced have also since been fixed. The ranking-quality question itself has not been re-measured since those fixes and remains WITHDRAWN pending a new preregistered measurement; see also ranking_vs_volume, a separate, later finding that VSigma's own ordering (which replaced this composite as the shipped default) is indistinguishable from a volume sort rather than actively harmful.","safe_to_say":"A blinded evaluation of our full scored output against a plain volume ranking of the same data found the scored output performed worse, not better. We do not claim our ranking helps a reader find what matters faster than a volume sort would, and several specific defects the evaluation surfaced have since been fixed but not re-measured."},{"id":"evidence_specific_attribution_precision","claim":"Naming the specific evidence checked for a 'not reached' holding (which channels were checked and found not to name it, versus a generic refusal sentence) reduces false-attribution claims without a material recall cost.","status":"UNPROVEN","figure":"Run 5 (n=96, random draw): +10.0 precision points, p=0.058. Run 6 (n=180, random draw): OLD made zero false positives across the whole sample, no discrimination possible. Run 7 (n=120, six predefined failure classes read directly off the resolver's own code, specifically built to find an OLD failure if one exists): precision delta -0.0043 (95% CI [-0.0128, +0.0042], p=1.000), FP/read delta -0.0417 (p=0.938); both arms ~97-98% precision even on cases engineered to be hard.","test":"Three independent blinded-reader confirmatory runs, increasingly adversarial sampling, zero case overlap across runs; a mandatory scorer audit before each result was trusted (Run 7's audit found and fixed a real classifier gap before scoring).","source":"research/attribution/RUN7_RESULTS.md","safe_to_say":"We tested whether naming the specific evidence checked, rather than a generic refusal sentence, measurably reduces false attribution claims. Three runs, the third built specifically to find a failure if one exists, found none: OLD's plain refusal sentence is already hard to fool. We keep the more specific wording because it is more auditable by construction, not because it is proven to be more precise."},{"id":"book_collapse","claim":"A book's 'effective bets' figure describes how many independent documented exposures it holds, and is more informative than the largest-single-bloc proxy it replaces.","status":"EXPLORATORY","figure":"participation ratio of the book's exposure matrix; on the shipped sample portfolio 2.82 against the old proxy's 2.00, and on a five-name technology book 1.37 against 1.00. Two of five test books differ by >= 0.5 bets. CORRECTED 10 September 2026 from 2.88/2.20 and 1.55/1.00: the original was measured with a bare PortfolioAnalyzer() and so on a poorer exposure graph than the product has, and the figure was additionally inflated by the TICKER: self-loop defect fixed the same day. Both preregistered conditions still pass on the corrected numbers, and the verdict remains stable across RANK_DECAY 0.5-0.8.","test":"28 unit tests fixing the behaviour at both ends (identical positions -> 1.0, disjoint -> n, hedge -> one axis, unmapped -> refuses), plus research/collapse_experiment.py run against the live resolver with three preregistered success conditions. Condition (c) FAILED on the first run at 1.78 vs 1.00 on a single-bet book, which forced the rank-decay correction; it passes at 1.18 and the verdict is stable across RANK_DECAY 0.5-0.8.","source":"research/SIGMALENS_SIGNATURE_CONCEPT.md","safe_to_say":"We report how many independent documented exposures a book holds, built from a curated exposure graph and position weights, with no price or return data. It is a structural description, not a forecast, and it has not yet been put in front of blinded readers, so we do not claim it helps anyone decide faster or better."},{"id":"keystone_evidence_load","claim":"A book's visible concentration can be traced to the citations underwriting it, and the share resting on any single citation can be measured.","status":"EXPLORATORY","figure":"on four of seven realistic books the keystone load is 0.76 to 1.00 (two of them at 1.00) and the keystone is the curated entity-to-ticker map; on the energy book every citation scores 0.00 because each holding carries 26 to 37 citations, and the module names nothing. Across the graph, one curated artefact underwrites 18 of 23 reachable instruments. The upper bound read '0.80 to 1.00' before 10 September 2026; KEYSTONE consumes the effective-bets figure and moved with the COLLAPSE correction of that date.","test":"16 unit tests, including a decisive pair of books that are structurally identical (1.00 effective bets each, indistinguishable to COLLAPSE) and evidentially opposite (load 1.00 vs 0.00), plus a 15-book adversarial battery and a sensitivity run on whether generic citation labels are treated as shared, which changed no result.","source":"research/SIGMALENS_NEXT_PRIMITIVE.md","safe_to_say":"We can say how much of the concentration we are able to see in a book depends on one citation, and name it. That measures our own visibility and our dependence on a source, not whether the source is wrong, and it says nothing about returns. It has not been put in front of blinded readers, so we do not claim it changes anyone's decision."},{"id":"portfolio_topology_bridge","claim":"A book's documented exposure forms separable regions, and the individual holdings that join them can be identified.","status":"EXPLORATORY","figure":"GLD is the sole structural bridge in ONE of seven realistic books (a long/short book of XOM, USO, GLD and AAPL, where removing gold separates the oil names from AAPL). It is 20% of that book and moves the effective-bets figure barely at all. Three of the seven return NO_UNIQUE_BRIDGE and two refuse. CORRECTED 10 September 2026 from 'three of seven': the original was measured with a bare PortfolioAnalyzer(), which has no storage, no profiler and a strictly poorer exposure graph than the API constructs. Re-measured against the production analyzer, the count is one.","test":"18 synthetic books forcing hub, bridge and redundant connector apart, plus order- and scale-invariance checks, plus seven real books reaching all five states. The obvious alternative construction (change in effective bets on removal) was implemented and FAILED: on a book where the bridge is known by construction it ranks the true bridge equal with two non-bridges. A defect found in production on 10 September 2026 is now a named regression test: the live resolver tags each holding with TICKER:<its own ticker>, a self-loop that survived the universal-entity strip as the sole residue of otherwise identical holdings, and an all-oil XOM/CVX/XLE/USO book was described on the product surface as '3 separate regions with nothing joining them' while COLLAPSE called it 1.2 bets on the same screen. It now refuses.","source":"research/topology/BRIDGE_RESULTS.md","safe_to_say":"We can show which regions a book's documented exposure separates into and which holding is standing in more than one of them. It uses no prices and no returns, naming a bridge is not advice to trade it, and it describes our curated graph rather than the world: an exposure we have not documented cannot appear. It has not been tested on human readers."},{"id":"event_structural_span","claim":"An event can be shown to tie together parts of a book that are otherwise structurally separate.","status":"UNPROVEN","figure":"the cross-region state fired ZERO times across 300 stored events against four realistic books; 110 event-book pairs resolved and every one of them was contained to a single region. Re-measured 10 September 2026 against the production analyzer (the original run used a bare PortfolioAnalyzer() with the profiler switched off, and resolved 60 pairs on a poorer graph). The verdict survived the correction unchanged. An intermediate run on the production graph appeared to fire the state 48 times, ALL of them in one all-oil book and all of them artefacts of the TICKER: self-loop defect fixed the same day; every one described an XOM/CVX/XLE/USO book hit by an oil shock as spanning three separate regions, which is the most concentrated outcome that book has.","test":"14 adversarial synthetic cases, all behaving, including the inversion the measure exists for: four holdings and 80% of book inside one region is structurally NARROWER than two holdings and 40% across two regions. Injecting a hypothetical US-China-shaped event into the real graph fires the state correctly, so the mechanism is sound and the case is absent from the feed rather than the measure being broken.","source":"research/structural_shock/STRUCTURAL_SHOCK_RESULTS.md","safe_to_say":"We built the ability to say whether an event touches several structurally separate parts of a book, and on our current event feed no event has ever done so. The capability is not in the product and we make no claim for it. Real events in the window tested were single-region stories and the measure correctly said so."},{"id":"event_mechanism_decomposition","claim":"An event's reach into a book can be decomposed by the documented mechanism it arrived through, and a reach that looks broad is often one mechanism repeated.","status":"WITHDRAWN","withdrawn_figures":["19 of 65 SEVERAL","19 (29%)","2 mechanisms","3 separate mechanisms","9 of 9 agreement","19 of 19 mix vocabularies"],"why_withdrawn":"The decomposition compared values that are not the same kind of object: sector classifications, curated risk categories and economic mechanisms were counted against each other as though they were comparable. Re-deciding all 65 real cases with semantic typing returns zero SEVERAL, so every 'several separate mechanisms' statement the product showed was unsupported - 14 of 19 were one mechanism plus gold's safe-haven exposure, which is attached to every entity in the map, and the other 5 could not be determined at all.","why_withdrawn_detail":"11 September 2026, after a hostile audit of where the channel strings come from. The decomposition compared values that are not the same kind of object. Of 232 channel rows across the 65 real multi-holding reaches, 137 are SECTOR classifications, 35 are curated RISK CATEGORIES, 43 record no channel at all, and only 16 are economic mechanisms - every one of those 16 being GLD's 'safe-haven demand', which commodity_fund_exposures attaches to EVERY key of ENTITY_MAP at a measured entity-map share of 1.00. Re-deciding all 65 cases with semantic typing gives 60 SINGLE_MECHANISM, 5 INSUFFICIENT and ZERO SEVERAL: of the 19 cases that shipped a 'several separate mechanisms' statement, 14 were one mechanism plus gold and 5 could not be determined. SPECIFIC FAILURES ON THE RECORD: (1) a gold standard built from ExposureGraph paths carrying channel=None reported 9 of 9 agreement, which was a null group matching a single-mechanism answer and not validation; (2) a provenance classifier consulting only ENTITY_MAP labelled curated commodity-fund exposures as derived and produced a false '19 of 19 mix vocabularies'; (3) independent gold labels exist for 2 of 65 cases because COUNTRY_CHANNEL does not key on the COMMODITY: and TICKER: codes the feed carries, and 54 channel=None paths had to be rejected; (4) the one gold case that disagreed is a SEVERAL, where the graph assigns all four holdings a single crude_supply channel. WHAT SURVIVES, and is still shown: the single-route statement, 46 of 65 cases, unchanged by typing and agreeing with the one gold case available to it. That is a statement about route identity, not about mechanism diversity, and it is not this claim. DOMINATED remains unobserved at its frozen 0.80 threshold and is withdrawn with the rest. See research/conduit_full_audit.md.","figure":"over 300 stored events against four realistic books, 65 event-book pairs reach two or more holdings; 46 of them (71%) arrive through a SINGLE mechanism and 19 (29%) through several. The two measures that could otherwise say this are constants on the same data: event_collapse reports 1.0 blocks in all 65 and splits_the_book is False in all 65.","test":"16 unit tests and a 22-case adversarial battery with no mismatches, including a deciding pair with identical reach count and identical reached weight that must produce opposite answers. Fund look-through normalisation is load-bearing and measured: without it an all-oil book hit by an oil shock reports two mechanisms, the most concentrated case there is described as diversified. Re-measured with SIGMALENS_VENDOR_PROFILES=off, where it degrades into naming the holdings it cannot place rather than into a wrong answer. The preregistered DOMINATED state fired ZERO times at its frozen 0.80 threshold and the threshold was NOT lowered to rescue it.","source":"research/information_gap/CONDUIT_RESULTS.md","safe_to_say":"We used to tell a reader when one event reached their holdings through several different mechanisms. We withdrew that. An audit of where those mechanism labels come from found we were comparing sector classifications against risk categories as though they were the same kind of thing, and that the only genuine mechanism in the data is attached to every entity in our map, so any book holding gold acquired a spurious second mechanism. We make no claim about mechanism diversity. We still say when everything an event reaches arrives by the same documented route, which is a statement about our own routing and not about the world, and it has never been tested on a human reader."}],"note":"A row with no safe_to_say may not be quoted anywhere. Withdrawn figures may appear only inside their own retraction."}