At a glance
- A vendor KPI that is not tied to a threshold and a defined action is a vanity metric. The point of a number is the decision it forces.
- Build the vendor scorecard from the risk-based logic the regulators already expect: identify the critical data and processes, then measure the vendors that touch them.
- Every metric needs four things: a definition, a threshold, an owner, and the action a breach triggers.
- Metrics are oversight evidence. Collected and never reviewed, they catch nothing and prove nothing.
- Centralized, aggregated review of vendor metrics catches drift earlier than periodic check-ins, and costs less attention.
This is the measurement layer of vendor management in clinical trials, and it is how vendor oversight turns from activity into signal.
Metrics, KPIs, and thresholds are not the same
Three terms get used loosely, and the looseness is why so many vendor dashboards catch nothing. A metric measures whether something is performing: query rate, data-entry timeliness, sample turnaround. A KPI is the small subset of metrics tied most directly to what matters for the trial, the ones you would escalate on. A threshold is the line that turns a KPI from a number into a trigger: the value at which “keep watching” becomes “act.”
The discipline that matters is the last one. A scorecard full of green-and-red tiles is decoration unless each tile has a pre-agreed threshold and a defined response. Practitioners often borrow the language of quality tolerance limits and key risk indicators here, and that framing is useful, but the underlying requirement is simpler and older: decide in advance what level of performance is unacceptable, and decide what happens when a vendor crosses it.
Build the scorecard from risk, not from what’s easy to count
The temptation is to measure whatever the system reports. The better starting point is the one the regulators already prescribe for monitoring. FDA’s risk-based monitoring guidance directs sponsors, at the protocol design stage, to identify the critical data and processes necessary for human subject protection and maintaining data integrity, and then to perform a risk assessment on them. Its 2023 question-and-answer companion adds that this risk assessment should weigh the potential causes, the likelihood of detection, and the severity of the consequences of each risk. That is exactly the logic a vendor scorecard should inherit: measure the vendors that touch your critical data and processes, on the dimensions where their failure would hurt, rather than on whatever is easiest to chart.
ICH E6(R3) ties this to the vendor relationship directly through the monitoring plan, which it says should be tailored to the identified risks to participant safety and the reliability of the results, and through its general rule that oversight should be fit for purpose and tailored to the trial’s risks. A vendor scorecard is one expression of that tailoring: the high-risk data vendor gets a richer, more frequent metric set than the low-risk courier, because the consequences of its drift are larger.
Every metric needs four things
A metric earns its place on the scorecard only if it carries all four of these:
- A definition. What is measured, how, and from what source, written down so the number means the same thing every month and to every reader.
- A threshold. The value that separates acceptable from not, set in advance. Without it, every review is a fresh argument about whether a number is “bad.”
- An owner. A named person who reviews it and is accountable for acting, not “the team.”
- An action on breach. What happens when the threshold is crossed: a conversation, a CAPA request, an escalation, an audit trigger. The action is the whole reason the metric exists.
Run that test across a typical dashboard and most tiles fail it. Cutting the metrics that have no threshold or no action is not losing information; it is removing noise that was hiding the few signals that matter.
Example metrics, by what they protect
Good vendor KPIs map to the consequence the vendor could cause, which differs by vendor type. A data-management or EDC vendor is measured on data-quality signals: query rate and ageing, time to database lock, the volume of post-entry data changes. A central lab is measured on sample turnaround, result-reporting timeliness, and reconciliation discrepancies. A CRO running monitoring is measured on visit timeliness, action-item closure, and protocol-deviation trends. A drug-supply or logistics vendor is measured on on-time delivery, temperature excursions, and stock-out risk. In each case the metric is chosen because a bad value points at a real consequence for safety or data, and each is paired with the threshold and action that make it actionable.
To see why the threshold is the whole game, take query rate on an EDC vendor. The raw number (“312 open queries”) means nothing on its own. Tied to a definition (open queries per 100 data points), a threshold (escalate above an agreed rate), an owner (the data lead), and an action (a CAPA request if the rate stays high for two review cycles), the same number becomes a decision: cross the line twice and the vendor enters a defined remediation path rather than a vague conversation. The metric did its job not by existing but by forcing a response at a pre-agreed point.
The time to fix all of this is before the work starts, not after a problem. Agreeing the KPI set, the definitions, the thresholds, and the reporting cadence in the contract or quality agreement does two things: it gives the vendor a fair, known target rather than a moving one, and it secures your right to the underlying data the metrics depend on. Metrics introduced after a vendor is already underperforming arrive late and read as punitive. Metrics agreed at contracting are simply the terms of the relationship, and a good vendor builds its own reporting to meet them.
Review cadence and centralized monitoring
A metric reviewed too late is a metric that catches an incident after it has happened. The cadence should match the vendor’s risk: high-risk vendors reviewed monthly or more, lower-risk ones quarterly. The bigger leverage, though, is reviewing metrics centrally and in aggregate rather than vendor by vendor in isolation. FDA’s guidance is explicit that centralized monitoring, a systematic analytical evaluation of study conduct across sites carried out by sponsor personnel, helps sponsors aggregate and compare data and detect potential anomalies more quickly, and that focusing monitoring on the most critical data elements lets sponsors achieve quality without frequent routine visits everywhere. Applied to vendors, the same idea means watching the metric trends across vendors and over time, where a slow drift shows up well before a single month’s number trips a threshold.
Make the metrics oversight evidence
The final discipline is to treat metrics as part of the oversight record, not a transient dashboard. When a threshold is breached, the breach, the decision it triggered, and the outcome should be captured, because that trail is what demonstrates, at inspection, that your metrics were not just collected but acted on. A vendor scorecard that shows a red KPI three months running with no recorded response is worse than no scorecard, because it documents that you saw the problem and did nothing. The value of measurement is realised only in the decisions it drives and the record of those decisions.
Where vendor metrics go wrong
- Vanity metrics. Numbers with no threshold and no action. They fill a dashboard and catch nothing.
- Measuring the easy thing. Tracking what the system reports instead of what protects critical data and participants.
- Collected, never reviewed. Metrics gathered on a cadence nobody keeps, so a breach sits unseen.
- No owner. A KPI everyone can see and no one owns is a KPI no one acts on.
Where VendorVigilance fits. Vendor metrics only work when the definition, threshold, owner, and action live together and the breaches are recorded. VendorVigilance holds KPIs and deliverables in its Governance module, tracks performance against agreed targets, and surfaces it through fifteen built-in reports with filtering and export, all attached to each vendor in the central registry on a 21 CFR Part 11-compliant audit trail. So a red metric is not a tile on a screen that disappears next month; it is a recorded signal that triggers an action and stays in the oversight record. Explore the product.
The bottom line
Measure vendors on what protects the trial, give every metric a threshold, an owner, and an action, review them on a cadence set by risk and in aggregate so drift shows early, and keep the breaches and the responses as oversight evidence. A scorecard built that way is not reporting. It is an early-warning system that turns a vendor problem into a number, and a number into a decision, before it becomes an incident.
Sources
Dejan Murko
Dejan is the co-founder of Mayet, building software for biotech and pharma teams.
