"Our risk score dropped from 72 to 64." The sentence sounds like information. Whether it is depends entirely on what is inside the number — and in most committees, nobody in the room can say.
A cyber risk score is an index that summarises many signals — vulnerabilities, exposure, control status, threats, business context — on a single scale, such as 0 to 100 or A to F. It is not a unit of loss. It is a relative indicator: it compares over time and across business units, assets or suppliers. Used that way, it is one of the most practical tools in risk management. Used as if it were an absolute measure, it misleads.
Two families of score
The market calls very different things a "risk score". Separating the two families avoids most unproductive debates:
| Outside-in | Inside-out | |
|---|---|---|
| Source | Signals observable from the internet: exposed services, certificates, e-mail configuration, leaked credentials | Inventory, scanners, control status, business criticality |
| Typical use | Assessing and monitoring suppliers; industry comparison | Prioritising treatment; reporting internal posture |
| Limit | Cannot see internal controls or know what is critical to the business | Only as good as the inventory and the data quality |
The outside-in view is especially useful for third-party risk, where you have no access to the supplier's environment. But it does not replace assessment: a good external score says the shop front is tidy, not that the process is secure.
How a good score is built
Inside, a useful internal score combines four components:
- Exposure. What an attacker can reach — from the internet, from a partner, from inside the network.
- Weakness. Vulnerabilities and misconfigurations, weighted by severity (CVSS) and by likelihood of exploitation (EPSS, the KEV catalog).
- Control effectiveness. Tested, not declared. A control that exists in policy and was never verified should not improve the score.
- Business impact. The criticality of the asset or of the process it supports.
Then come four construction decisions that must be written down: how each signal is normalised to a common scale; what the weights are and who set them; how the score aggregates from asset to business unit and from unit to company; and how to walk back — from the number to the asset and the flaw that explain it.
Aggregation: where most scores go wrong
Picture a business unit with a hundred servers. Ninety-nine score 95. One of them — the internet-facing ERP, with an open vulnerability from the KEV catalog — scores 10. The unit's average is 94, and the dashboard shows green.
The attacker, however, only needs server number 100. The average is the worst aggregation choice for risk, because it dilutes exactly the asset that matters most. Better alternatives:
- Floor rule. If any critical asset has open confirmed exploitation, the unit cannot score above a defined cap, regardless of everything else.
- Weighted worst case. The unit's score gives more weight to the assets with the worst score and highest criticality.
- Distribution instead of a number. Showing how many critical assets sit in each band says more than any average.
Five signs the score misleads
- Nobody can explain why it changed. If the change cannot be traced to a cause, it does not guide action.
- It rises with activity, not with risk reduction. Published documents and completed trainings improve the score while exposure stays the same. It is the classic effect of turning a measure into a target: people start optimising the number.
- It uses an average with no floor. See the previous section.
- It has false precision. A score of 72.4 suggests an accuracy the input data does not have.
- It changes no decision. If nothing happens when the score drops, it is dashboard decoration.
Score or quantification?
The two approaches are complementary, and each answers a different kind of question better:
| Question | Best-suited tool |
|---|---|
| Where do we act this week? | Internal score, traceable to the asset and the flaw |
| Are we improving? | Score trend, with a stable methodology |
| Is this supplier acceptable? | External score plus supplier assessment |
| How much to invest? Does the control pay off? What insurance limit? | Quantification (CRQ) |
A score does not turn into money by multiplication. Multiplying a score by revenue or by number of records is not quantifying risk: it is applying arithmetic to a scale that was not built for it. When the question is in money, use a loss model.
Some teams replace the score with a decision tree, such as the SSVC published by CISA for vulnerabilities: instead of a number, the output is an action ("track", "act"). It is a useful alternative when the real purpose of the score was to decide treatment deadlines.
Questions to ask of any score
- What data goes in, where does it come from and how often is it updated?
- What are the weights, and who set them?
- How does it aggregate — average, worst case, floor?
- Can I open the number down to the asset and the flaw that explain it?
- What do we actually do when it drops?
A score that answers all five well is a management tool. One that does not is a pretty number.
Sources consulted on 4 October 2026: NIST IR 8286B, Prioritizing Cybersecurity Risk for Enterprise Risk Management; Exploit Prediction Scoring System (EPSS), maintained by FIRST; CISA's Known Exploited Vulnerabilities catalog and Stakeholder-Specific Vulnerability Categorization (SSVC). Numerical examples are illustrative. This article describes principles for building indicators and does not evaluate rating products or vendors.