Scoring methodology.
How risk scores combine static security posture with threat intelligence, exploits in the wild, sector susceptibility and dark-web data, cross-input correlation, severity mapping and asset-weighted aggregation.
Every scoring function, weight and signal is documented. If the model is wrong for your organisation you should be able to refute it with evidence. Black-box ratings are theatre.
Four input streams
The score is a weighted combination of four independent streams. Each is collected from public sources, requires no access to your systems, and is independently auditable.
| Stream A · 45%: Static posture | Findings from the six scanning modules: TLS configuration, HTTP security headers, exposed services, DNS hygiene, certificate management and cloud posture. Snapshotted continuously. |
| Stream B · 25%: Threat intelligence | CVE feed cross-referenced against detected versions, known-exploited vulnerabilities (CISA KEV), exploit availability scores (EPSS), and active-campaign sightings from sector-peer telemetry. |
| Stream C · 20%: Sector susceptibility | Industry attack base rates, geographic threat density, vertical-specific ransomware activity, and breach frequency at companies of comparable size and posture. |
| Stream D · 10%: Dark-web signal | Credentials, source code and customer data observed in criminal marketplaces, ransom-leak sites, paste sites and underground forums attributable to your organisation. |
Severity is contextual
A CVSS base score alone is not enough. Severity is base × exploitability × exposure × asset criticality, the same vulnerability on a marketing microsite and a customer billing API does not contribute the same risk.
| Critical, active exploitation observed | CVSS 9.0+, KEV listed, EPSS above 0.7. Pre-auth RCE, mass scanning, or compromise tooling published. Carries a 4× multiplier into the module score. |
| High, reliable exploitation path | CVSS 7.0–8.9, EPSS 0.3–0.7. Auth bypass, SQL injection, privilege escalation. Public proof-of-concept exists but no mass exploitation yet. 2.5× multiplier. |
| Medium, defensible misconfiguration | CVSS 4.0–6.9. Weak TLS, missing security headers, information disclosure, deprecated services. 1.5× multiplier, and aggregates fast across a large estate. |
| Low, hygiene signal | CVSS 0.1–3.9. Server banners, version disclosure, stylistic issues. 1× weight; informational on the report but still visible. |
Asset-weighted, geometric mean
Module scores aggregate into the company rating with asset weighting, so a single critical exposure on a business-critical asset is not averaged away by a large estate of healthy ones.
- Findings roll into module scores; module scores roll into one company rating
- Asset criticality weights the contribution of each finding
- A geometric mean prevents a long tail of healthy assets masking a severe one
- The rubric is versioned, so a score is reproducible against a given version
Refute any finding
Any finding can be disputed from the dashboard with a reason and supporting evidence.
- Reasons: wrong attribution, mitigated, accepted risk, false positive
- An analyst reviews every dispute
- Accepted disputes update the score on the next scan
- The decision and its rationale are recorded in the audit trail
A score nobody can interrogate is a score nobody will act on
Security ratings have a credibility problem, and it is largely self-inflicted. Vendors treat the model as proprietary, rated companies cannot understand why their number moved, and the first response to an uncomfortable score is to attack the method rather than fix the finding. Everyone loses that argument.
So the model is published. Not a summary of the philosophy: the actual weighting, how severity is applied, how findings aggregate across assets, and how versioning works. A buyer should be able to evaluate whether they believe the model before they trust the number it produces.
The practical test is a supplier dispute. When you send a vendor findings about their estate, you need a method that survives their security team reading it carefully. A proprietary black box does not survive that; a published deterministic model does.
The same evidence must always produce the same score
Determinism is the property that makes trending possible. If a score can move because the model was adjusted, then quarter-on-quarter comparison is meaningless, you cannot tell whether posture improved or scoring got more lenient.
So the model is versioned. A change in methodology produces a new version, and historical scores remain attributable to the version that produced them. When a score moves, the cause is unambiguous: the estate changed, not our opinion of it.
This is the difference between a rating that can be governed and one that can only be observed. A board asking whether two years of investment moved anything needs the second answer to be impossible.
How to argue with a finding, and win
Every finding carries the evidence it was derived from: the record, the response, the certificate, the log entry. That is deliberate: it makes disputes factual rather than rhetorical.
There are three ways a finding legitimately falls. The evidence is stale and the condition no longer holds. The attribution is wrong and the asset is not yours. Or the observation is technically correct but the inference does not follow: a version banner that reports vulnerable because a distribution maintainer backported the patch without changing the string, which is the most common genuine false positive in this category.
All three are checkable, and a finding that fails any of them should go. A platform that cannot dispose of wrong findings accumulates them until practitioners stop reading the output, at which point the rating is decorative.
What the model deliberately does not attempt to measure
The model scores what is externally observable. It does not attempt to infer internal controls from external signals, and that restraint is deliberate rather than a limitation awaiting a future release.
The temptation in this category is to imply that a strong external posture predicts strong internal governance. There is a correlation, an organisation that leaves certificates to expire is not usually one running rigorous access reviews, but it is a correlation, not a measurement, and presenting it as the latter is how ratings platforms lose the trust of the practitioners who have to act on them.
So segmentation, backup testing, access governance, incident response and training are simply absent from the model. Those need an audit, and the rating's job is to tell you where audit capacity is best spent.
Common questions
Is the scoring model published in full?
Yes, weighting, severity handling, aggregation and versioning. A model a buyer cannot evaluate is a model whose output they will not defend in a supplier dispute.
Why does versioning matter?
Because it makes trending honest. If a score can move when we adjust the model, quarter-on-quarter comparison tells you nothing. Historical scores stay attributable to the version that produced them.
How do we dispute a finding?
With the evidence it carries. A finding falls if the evidence is stale, the attribution is wrong, or the inference does not follow, the last usually being a backported patch that leaves a version string looking vulnerable.
Do you infer internal controls from external signals?
No, deliberately. There is a correlation between external hygiene and internal rigour, but it is a correlation rather than a measurement, and presenting it as the latter is how ratings platforms lose practitioner trust.
Can two assessments of the same company differ?
Not on the same evidence and model version: that is what deterministic means. They can differ if the estate changed between assessments, which is the point of continuous rating.
Run the model against your own company
The quickest way to test a scoring model is on a company whose posture you already know.
Free for your own organisation. No access required.