Summary
- OpenSSF Scorecard assigns 0–10 results to selected automated checks and produces a risk-weighted aggregate; it says the checks are heuristics, with false positives and false negatives.
- The project explicitly rejects the idea that Scorecard is a definitive, one-size-fits-all report or requirement. An aggregate can conceal different underlying check combinations.
- A consumer can use a scan as evidence, but the scan does not identify the dependency release in scope, set the consumer's risk tolerance, validate every control, or choose accept, monitor, replace or remediate.
A number with an intentionally limited job
Security decisions often begin with an awkward surplus of signals. A repository may have workflow files, tags, issues, release assets, configuration, contributors and an assortment of public security claims. A person deciding whether to depend on it cannot inspect every clue from scratch each time. A score has value because it compresses selected observations into something that can be compared and investigated.
OpenSSF Scorecard is explicit about that job. It automates assessments of selected security-related heuristics, gives individual checks a result from zero to ten and calculates an aggregate using risk weights. It is designed to help maintainers improve practice and help consumers think about dependency risk. It does not say that a high number is a release approval, a complete security assessment, a contractual representation or a decision already taken by the consumer.
That restraint matters. The project calls every part of its construction opinionated: which checks appear, which are omitted, how much each matters and how the output is calculated. It also warns that checks can produce false positives and false negatives. Its stated non-goals include being a definitive report or a requirement that all projects must satisfy. An aggregate, it says, does not reveal the individual behaviours that produced it; different combinations can lead to the same number, and the result can change as heuristics evolve.
Those are not defects accidentally admitted in small print. They describe the boundary that makes automation honest. A Scorecard result is a measured view, not the whole repository; still less is it the whole software supply chain around a consumer's use of that repository.
The aggregate is a map, not the territory
The temptation is strongest when the score is convenient. A dashboard can display one aggregate beside hundreds of repository names. A badge can make the number portable. A procurement or engineering discussion can then be shortened to a threshold: above, below, pass, fail.
But the aggregate is a weighted average, not an explanation. The Scorecard documentation gives checks different risk levels, and the README describes corresponding weights. Two targets can therefore reach the same aggregate through unlike conditions. One may have a strong result on a control relevant to a particular consumer and a weak result elsewhere; another may reverse that pattern. The number alone does not say which difference matters to a package deployed in a particular architecture.
The observation surface also has edges. Public weekly scans are a defined programme, and the REST API has documented omissions for some checks at scale. A command run against a target can have different access and evidence than an API result. Checks can be inconclusive. The visible reference, the time of the scan, the tool version, accessible metadata and selected check logic all help explain what a result means.
That is why a dashboard should not silently become a historical claim. Seeing a score today does not prove the score for a prior release, the state of a private branch, the content of an unreached service, the security of a distributed artifact or the experience of a downstream operator. It proves only that the described tool observed the described target through its available methods at a particular point.
A check observes its own question
Individual checks are more informative than a single aggregate because they name what they attempted to see. They are still not universal verdicts. The Maintained documentation, for example, uses observable activity over a defined period, while also recognizing that software with little visible activity is not necessarily unsafe or abandoned. The SBOM check looks for a bill of materials in specified source, pipeline or release locations. Finding one is useful evidence of presence; it is not proof that the inventory is complete, current, accurate for a particular deployed binary or sufficient for a consumer's exposure analysis.
The same discipline applies to a low result. The documentation repeatedly explains that a practice may exist without being detected and that a low score is not a definitive indication of risk. A failed automated observation is a reason to inspect a stated condition. It is not proof of negligence, compromise, a vulnerable product or a rejected dependency.
Scorecard accommodates this gap rather than pretending it does not exist. Its maintainer-annotation mechanism lets a project add limited context such as test-data, remediated, not-applicable, not-supported or not-detected. The annotation is valuable because it records a claim that the generic probe may not tell the entire story. It remains a maintainer-supplied explanation, not independent validation and not an instruction to a consumer.
This creates three distinct records. First comes the tool's observation: target, time, version, checks, reasons and results. Second comes a maintainer's contextual statement where one is supplied. Third comes a consuming organisation's decision about a concrete dependency and use. None can honestly stand in for the other two.
The decision begins where the scan stops
A dependency decision has a different object from a repository scan. The consumer may rely on a packaged release, a pinned commit, a transitive component, an internal mirror or a build output. It may expose the component to different data, privileges, networks and recovery obligations than another consumer. It may have compensating controls, an urgent compatibility need, an alternative supplier, a sunset plan or a legal constraint. Scorecard cannot select among those facts because they belong to the consumer and its operating context.
Nor can a score determine what action follows. A team can accept a dependency with a documented exception; defer a decision while asking for evidence; isolate a component; prefer an alternative; add monitoring; require a later recheck; or decide not to use it. Each action has a decision-maker, rationale and expiry condition. Those are governance choices, not outputs of a weighted average.
This is the useful line to preserve: Scorecard can make a question sharper without answering it by authority. A low Code-Review observation does not prove that a particular change was unreviewed. A high Signed-Releases observation does not prove that the release being evaluated is genuine or appropriate. A discovered security-policy file does not prove an incident response succeeded. Automation yields bounded evidence. Responsibility for reliance stays with the actor that will bear the consequence.
Keep a scan-to-dependency decision record
The repair is modest. A consequential dependency file should retain the repository identifier and resolved reference, the Scorecard version and scan time, accessible-data limits, the aggregate and the specific checks that mattered. It should preserve reasons and details, along with any maintainer annotation and its source. That makes the measurement reproducible enough to be challenged without turning a public score into a secret gate.
The second half must be owned by the consumer. It should name the package, release or commit actually under consideration; the relevant environment and exposure; independent evidence; accountable decision-maker; threshold or exception rationale; action selected; and next review or trigger. If the decision is an exception, the record should say when it expires. If it relies on a compensating control, the owner and test condition should be visible.
The record does not demand that every project adopt one security regime. It simply prevents a compact tool output from laundering a local judgment into an apparently objective verdict. The score remains useful because its boundary remains legible.
Sources
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
