Econometric research methodology
EDGAR Factor Lab
Separate broad-market sensitivity, sector-specific movement, filing change, and the price response after public SEC evidence arrives. Every estimate carries its sample, dates, model definition, coverage, and limitations.
Purpose
A diagnostic workspace, not a black-box score
The lab joins a provenance-aware market-price pipeline with source-linked SEC financial evidence. It does not collapse every result into one recommendation. Instead, it shows the separate quantities that support an interpretation and leaves unavailable values unavailable.
Return behavior
Regression estimates include uncertainty, overlap, and stability diagnostics.
Filing evidence
Period changes preserve reporting dates, coverage, and source accessions.
Reusable output
Versioned JSON gives researchers and AI assistants the same definitions.
Systematic sensitivity
Market beta with HAC uncertainty
The model converts positive price observations into close-to-close log returns and keeps only intervals whose start and end dates match the security and benchmark exactly. Prices and returns are not forward-filled across unmatched intervals. Version 1 estimates daily models only. The price transport excludes the current UTC calendar date so a still-forming Yahoo daily candle cannot enter a regression; the prior U.S. session becomes eligible after midnight UTC.
rᵢ,ₜ = α + βₘrₘ,ₜ + εₜandβₘ = Cov(rᵢ, rₘ) / Var(rₘ)Point estimate
Ordinary least squares estimates the intercept and market beta. The standard model requires at least 126 aligned return intervals and at least 80% coverage of eligible benchmark intervals. The output also reports correlation, R², adjusted R², residual volatility, influential dates, and the effective sample period.
Newey–West inference
Beta and intercept standard errors use a Newey–West heteroskedasticity- and autocorrelation-consistent covariance estimate. The automatic lag is the integer part of 4(n / 100)2/9, bounded by the available sample. The displayed 95% interval is β̂ ± 1.96 × HAC standard error.
Regime comparison
Upside and downside beta
The aligned sample is split by the sign of the benchmark return. One OLS market model is estimated when the benchmark return is positive and another when it is negative. Each side requires at least 30 observations; zero-return benchmark intervals belong to neither subset. Because sign-filtered observations are not consecutive trading sessions, each conditional regression uses a lag-0 HAC (HC1) covariance rather than serial-correlation lags.
β asymmetry = β downside − β upsideA positive difference means the estimated historical sensitivity was higher on negative benchmark sessions. This is a conditional sample comparison, not a crash forecast, tail-loss probability, or proof that the two coefficients differ statistically.
Orthogonalized exposure
Independent sector sensitivity
A sector benchmark often rises and falls with the broad market. Using it directly beside the market can blur the meaning of both coefficients. The lab first removes the sector benchmark's fitted broad-market component. The remaining sector residual represents movement not explained by that market model.
rₛ,ₜ = aₛ + bₛrₘ,ₜ + uₛ,ₜrᵢ,ₜ = α + βₘrₘ,ₜ + βₛ⊥uₛ,ₜ + εₜβₛ⊥ is sensitivity to the independent part of the selected sector proxy, after broad-market movement is removed. It is not the same as a simple company-versus-sector beta. Results retain the sector proxy, overlap, effective dates, and residual volatility so users can judge whether the comparison fits the company's business. The estimate is withheld unless at least 80% of eligible benchmark intervals have exact company, SPY, and sector-proxy matches.
Public-information event
Price response after a filing date
When a supported SEC acceptance timestamp is available, only a filing accepted before the regular 9:30 a.m. America/New_York open uses that session's ending return. An intraday, post-close, or non-trading-day acceptance starts with the next benchmark session. A date-only filing is also assigned strictly to the next session. Expected return comes from the broad-market and independent-sector model estimated over as many as 252 earlier aligned sessions. The estimation sample ends 20 sessions before the event and requires at least 180 observations.
ARₜ = rᵢ,ₜ − r̂ᵢ,ₜCARₕ = exp(Σ ARₜ) − 1, for h = 1, 5, or 20 sessionsStandardized responseₕ = Σ ARₜ / (σ residual × √h)The Filing–Market Map uses the 20-session standardized response. Standardization makes responses more comparable across securities with different residual volatility; it is not a p-value. The event output retains the acceptance timestamp, assigned return interval, and timing-quality code. Every published 1-, 5-, or 20-session window must contain each expected benchmark interval in all three price series; a suspension or missing quote withholds that horizon.
Filing–Market Map
SEC Filing Change z and EDGAR Evidence Gap
The map keeps reported change on one axis and model-adjusted price response on the other. Filing components are measured as period-over-period percentage-point changes, then normalized against other valid issuers in the same selected cohort and template. The focus issuer is excluded from its own peer reference. Peer reports are the latest point-in-time snapshots available when the lab runs, not a reconstructed cross-section frozen at the focus filing's event time; the API therefore reports peer filing-clock dispersion.
z = clip((company change − peer median) / (1.4826 × peer MAD), −3, +3)Each component requires at least eight valid other issuers. If median absolute deviation is zero, scale falls back to IQR / 1.349, then sample standard deviation. If dispersion is still zero, the component is unavailable. Raw changes are not winsorized; clipping the normalized z-score limits outlier influence.
Non-financial template
- Change in year-over-year revenue growth
- 25%
- Change in operating margin
- 25%
- Change in free-cash-flow margin
- 25%
- Change in book equity / assets
- 12.5%
- Change in cash / assets
- 12.5%
Bank and insurance template
Used for the credit-bank and insurance research cohorts.
- Change in year-over-year revenue growth
- 33⅓%
- Change in net margin
- 33⅓%
- Change in book equity / assets
- 33⅓%
Current and comparable-prior values must both exist. Missing values are never replaced with zero, and fixed weights are never silently redistributed. Version 1 requires every component in the applicable template; otherwise the composite Filing Change z is unavailable. Each change is recomputed from the current value minus its comparable-prior value. A separately supplied change is retained only as an audit consistency check.
Derived disagreement measure
EDGAR Evidence Gap
clip(Filing Change z, −3, +3) − clip(20-day standardized response, −3, +3)The range is −6 to +6. Positive means reported filing change was stronger than the model-adjusted price response; negative means price response was stronger than reported change. The map treats an absolute axis value below 0.5 as a neutral band.
Diagnostic only—not valuation, a forecast, a significance test, or a trade signal.Machine-readable research
One model definition for people and AI assistants
The versioned response identifies the security and benchmarks, effective dates, observation counts, overlap, return and price basis, model coefficients, confidence interval, diagnostics, filing-event windows, peer coverage, SEC evidence, calculation definitions, and warnings. Missing outputs remain explicit rather than being inferred.
Interpretation limits
What the lab does not establish
No EDGAR factor beta yet
Version 1 does not publish a historical high-minus-low EDGAR factor-return series and does not estimate βEDGAR. That requires point-in-time scores, historical membership, delisted securities, and effective-dated security mapping.
Current research universe
Peer normalization uses the available curated cohort at the calculation date. It is not an exhaustive industry index or a survivorship-free reconstruction of the investable market, and peer filings need not have become public on the same date.
Model uncertainty remains
HAC inference addresses certain residual variance and serial correlation patterns; it does not repair omitted variables, structural breaks, stale prices, non-synchronous trading, or an unsuitable benchmark.
Filing dates are imperfect event times
SEC acceptance timestamps govern session alignment when they are available, but daily returns cannot isolate the intraday response from earnings calls, guidance, macro news, or other information arriving during the event window.
Observed sessions are a proxy
Session sequences are inferred from observed SPY dates rather than a separately licensed official exchange calendar. Missing expected company or benchmark intervals withhold the affected event horizon.
Sector proxies are approximations
A traded sector benchmark can differ from a company's actual revenue mix, geography, capital structure, or economic peers. Independent sensitivity remains conditional on that choice.
Historical association is not causality
Betas, abnormal returns, peer z-scores, and the Evidence Gap are descriptive research outputs. None establishes mispricing, expected return, financial strength, or future performance.