Partial data¶
Part of Recommendations and Risk.
A scan step reaching completion is not the same as it succeeding. OSS IQ depends on four external sources — the package registry, GitHub, OSV and EPSS — and any of them can be firewalled, rate-limited or simply down. A report built on missing data must never be indistinguishable from a clean one.
What is tracked¶
Every fetch returns its outcome as a value alongside its data:
Status |
Meaning |
|---|---|
|
the source delivered |
|
some chunks failed; the rest of the data is real |
|
nothing came back — connection, timeout or HTTP failure |
|
the source’s quota was exhausted mid-scan |
Three of the seven scan steps report one:
Step |
Reports status? |
Why |
|---|---|---|
|
✅ |
four fetches — repo info, commits, activity, READMEs — combined into one status |
|
✅ |
one batch |
|
✅ |
one batch |
|
❌ |
per-package cached registry reads, not one batch run — deferred deliberately |
|
❌ |
no external data source; a status would be a category error |
Note
Two different combination rules, on purpose. Merging the fetches that make up one source treats
a mix of success and failure as partial — one failed GitHub stream out of four does not mean
GitHub was unreachable. Merging across all sources into the report-level verdict is a plain
worst-of — one failed source taints the whole report.
What breaks when each source degrades¶
This is the part that matters, because degradation does not produce errors. It produces quieter output:
Degraded |
Direct effect |
The recommendation you get |
|---|---|---|
OSV ( |
no CVEs on any record |
|
EPSS |
CVEs present, all unscored |
The strategy treats unscored as exploitable and moves more packages than usual; triage sees no scores, finds no exploit signal, and falls back to |
GitHub ( |
no maintenance state, no stability, weaker deprecation evidence |
No |
Registry ( |
not tracked |
A thin or failed registry read degrades the ladder itself, with no status to show for it. |
Where degradation is visible¶
Surface |
Shown as |
Present? |
|---|---|---|
Console stepper |
per-step icon and suffix — |
✅ |
Console, after the scan |
a stderr warning block: “Some data sources did not fully respond, so this report may be based on incomplete data”, listing each degraded step and what caused it — |
✅ |
Console, before the scan |
a stderr warning when GitHub’s remaining quota does not cover what the scan is about to need: |
✅ |
|
|
✅ |
|
|
✅ |
HTML report |
an Incomplete data banner above the table, naming each degraded source and carrying |
✅ |
Exit code |
non-zero (MCP: a titled error) when |
✅ opt out with |
Note
Every surface now says so, but they do not say it equally well. Machine-readable output
(--format agent, MCP, export) carries the degradation inside the document, where a consumer
cannot process the result without also being handed the caveat. The console warning and the HTML
banner sit beside the result and can be scrolled past. If you are automating on top of OSS IQ,
read data_completeness rather than trusting that someone saw the banner.
The governing rule¶
Unknown is not zero.
Both pipelines follow it, in opposite directions, and both are correct:
Triage is conservative about accusing. A package whose repository could not be measured is never marked for refactoring. Unknown is not unstable.
The strategy is conservative about writing. A CVE with no EPSS score counts as exploitable. Absent evidence must not silently read as absent risk when the output is a change to your manifest.
The difference is the consequence of being wrong. Triage produces advice a human reads; the strategy produces a version a tool writes.