September 2026 note: EPSS v4 began publishing on March 17, 2025. EPSS v5 superseded it on June 15, 2026. This article preserves the v4 transition because the governance lesson applies to every model change. Current implementation decisions should use the latest FIRST EPSS documentation and data.
When EPSS v4 went live, many vulnerability teams immediately compared yesterday’s score with today’s. It was a natural comparison, but a potentially misleading one.
A score can move because threat signals changed. It can also move because the model changed. Those are different events with different operational meaning. If a program cannot distinguish them, a model upgrade can silently rewrite patch queues, service-level targets, exceptions, and executive trends overnight.
The most important lesson from EPSS v4 was therefore not a particular feature or score distribution. It was that predictive security metrics require model governance.
What EPSS answers
The Exploit Prediction Scoring System publishes a daily probability between 0 and 1 for each scored CVE. It estimates the likelihood that exploitation activity will be observed in the next 30 days.
That question is operationally useful because technical severity and exploitation probability are not the same thing. CVSS describes vulnerability characteristics and potential impact; EPSS helps estimate likely attacker activity. Neither determines the business consequence of a vulnerable instance in your environment.
FIRST also publishes a percentile. Probability and percentile serve different purposes:
- Probability estimates the chance of observed exploitation activity in the next 30 days.
- Percentile shows where the CVE ranks relative to other scored vulnerabilities.
A high percentile can still accompany a modest absolute probability when most vulnerabilities have very low scores. Leaders should know which value a dashboard uses and why.
What changes at a model boundary
FIRST’s historical data guidance identifies the production boundaries for each model version. A time series crossing one of those dates contains a methodology change, not just new threat evidence.
That affects at least four parts of a vulnerability program.
1. Threshold populations
If a workflow labels EPSS scores above a fixed value as “urgent,” a new model may move many CVEs across that line. The program’s workload can change even when the asset population does not.
2. Trend reporting
An executive chart may show fewer high-probability vulnerabilities after a model transition. That is not evidence that the environment became safer. It may be a calibration effect.
3. Exceptions and automation
Automated ticket creation, due dates, or exception reviews may depend on a score threshold. A version change can alter those decisions without a human explicitly changing policy.
4. Historical comparisons
Comparing pre-transition and post-transition scores as one continuous series can create false narratives. Retain the model version and score date with every decision record.
A controlled migration pattern
When a scoring model changes, I treat it like a material rules-engine change:
- Freeze a reproducible baseline. Retain the old and new score datasets, calculation date, model version, and affected asset population.
- Run both models through current policy. Measure how many items change tier, owner, due date, or escalation path.
- Inspect consequential movements. Sample assets moving both up and down, especially exposed and high-impact services.
- Backtest the policy. Evaluate coverage, effort, and operational capacity rather than celebrating a smaller queue.
- Approve threshold changes explicitly. Do not let a model release become an unreviewed risk-appetite decision.
- Annotate reporting. Mark the transition date and explain the expected discontinuity to leadership.
- Monitor after cutover. Watch queue volume, aged exposure, emergency work, exceptions, and missed exploitation signals.
The executive conversation
An executive does not need a tour of the model’s feature engineering. They need to understand the decision impact:
- Did our definition of urgent change?
- How did remediation demand change by business service?
- Which previously accepted exposures are now above tolerance?
- Are we gaining better coverage, reducing effort, or merely moving a threshold?
- What remains outside the model—asset context, compensating controls, and consequence?
That is the bridge between data science and security leadership. The model provides evidence. Governance determines how that evidence changes action.
A durable control for the next version
Store these fields with every EPSS-informed prioritization decision:
- CVE identifier
- EPSS probability and percentile
- Score publication date
- EPSS model version
- Asset or service context
- Other decision signals, including KEV and exposure
- Resulting action, owner, and rationale
That record makes a decision explainable months later and makes model transitions testable instead of anecdotal.
Lessons learned
The v4 transition made one weakness easy to see: a feed can keep working while the control built on it changes underneath you. A successful ingestion would not have revealed movements in threshold populations, automated actions, historical comparisons, or remediation demand.
I still consider EPSS a useful public signal. I just do not treat it as a permanent severity label. Retain the model version, run candidate scores through the policy that consumes them, send material threshold changes to the accountable owner, and mark the reporting boundary. Before cutover, compare changes in tiers, owners, due dates, exceptions, capacity, and high-consequence assets; repeat that reconciliation after production changes.
At the next model boundary, the EPSS v5 production comparison measured substantial movement across fixed thresholds. Those thresholds belong inside a broader EPSS operating policy that keeps exposure and consequence alongside probability.
Does your executive reporting preserve the EPSS model version behind each trend, or could a model boundary still appear as an unexplained improvement or deterioration?
Sources and disclosures
This is an independent retrospective based on public FIRST documentation and the author’s vulnerability-management experience. It is not affiliated with or endorsed by FIRST. No employer, customer, private repository, or proprietary scoring implementation is described.



