A practical business guide

Vendor evaluation scorecard: a template with a worked example

Put evidence beside every score, separate requirements from preferences, and see whether your preferred vendor survives a change in assumptions.

Who it is for
Operations leaders and small teams choosing software or a recurring service.
What you will leave with
A documented vendor shortlist, weighted comparison, and next-step decision.
Published 6 min read
In this guide

The answer at a glance

A vendor evaluation scorecard compares qualified suppliers against the same criteria. First check mandatory requirements. Then define evidence-backed scores, assign weights totaling 100, calculate the results, and test whether reasonable changes in priorities alter the ranking. Keep missing evidence visible instead of assigning it an invented score.

Start with the purchase you actually need to make

Write one sentence that names the work, users, constraint, and decision owner. For example: “Choose a support platform for 12 agents that can import our ticket history and operate within an approved first-year budget.” This prevents a polished demo from changing the assignment halfway through evaluation. Record the current process as an alternative: keeping it may be reasonable when every proposal introduces more work than it removes.

Use the same evaluation packet for every vendor: representative tasks, sample data, expected outputs, pricing assumptions, and a deadline for clarifications. Agree what counts as proof. A feature mentioned on a pricing page, a successful trial using your test data, and a contractual commitment answer different questions. Attach a dated reference to each observation so a second reviewer can follow the reasoning.

Scotland’s Procurement Journey describes evaluation matrices as a way to score supplier responses against predefined criteria. It also emphasizes that results depend on the chosen criteria and weights, and that cost assessment should consider acquisition, operation, and end-of-life costs. The small-team worksheet here uses those general principles with its own scoring model.[1]

Apply mandatory requirements before awarding points

A mandatory requirement is a condition whose failure makes the purchase unsuitable. Give it a pass, fail, or pending status before the weighted comparison. Examples might include importing a required file format, an approved deployment region, or fitting a fixed budget ceiling. Each requirement needs a precise check and an owner. “Enterprise ready” is too vague to be a useful gate.

Keep the gate list short enough to defend. A preference becomes a gate only when the business cannot accept an alternative. If a supplier misses a gate, exclude it or explicitly revise the requirement for everyone before reconsidering the shortlist. A high score for attractive secondary features cannot compensate for a failed mandatory condition.

Pending means the evidence is incomplete. If a vendor has not demonstrated an export, record “pending: trial export required,” the person responsible, and a response date. Do not silently convert that uncertainty into a pass or a low capability score. A vendor that demonstrated a weak export and one that has not demonstrated any export pose different follow-up questions.

  • Define the requirement as an observable pass/fail condition.
  • Specify the evidence, evaluator, and deadline.
  • Resolve pending gates before approving the purchase.
  • Record rejected vendors and the requirement they failed.

Make a score mean the same thing for every vendor

Use a 1–5 scale for criteria that have enough evidence to judge. Write anchors for each criterion before viewing totals. For workflow fit, a 3 might mean completing all core tasks with documented manual workarounds; a 5 might mean completing the same tasks without those workarounds. A 4 represents a demonstrated improvement between those anchors. Missing evidence stays blank, with an open question attached.

For this worksheet, weighted points = weight × (score − 1) ÷ 4. A score of 1 receives none of that criterion’s points, 3 receives half, and 5 receives all. The maximum total is 100 when weights sum to 100. This normalization is a chosen convention, not a probability of success. Scores still depend on judgment, and equal gaps between scores are an assumption.

Separate criteria that measure different things. Workflow fit can measure task completion; implementation effort can measure migration and training. Do not award the same easy-onboarding benefit under five labels. For cost, define the user count, usage, term, support, migration, and exit assumptions. In the example, first-year cost bands are 5 at $12,000 or below, 4 above $12,000 to $15,000, 3 above $15,000 to $18,000, 2 above $18,000 to $21,000, and 1 above $21,000. These are fictional budget choices.

Worked example: three fictional support-platform vendors

The following vendors, observations, and prices are invented to demonstrate the method. All three have passed the mandatory checks. Vendor A completes the test workflow most cleanly; Vendor B costs less; Vendor C offers the easiest implementation. Assumed first-year totals are $16,800 for A, $11,400 for B, and $19,200 for C, producing cost scores of 3, 5, and 2 under the published bands.

Baseline comparison: 1–5 scores, with weights totaling 100
CriterionWeightVendor AVendor BVendor C
Workflow fit35544
First-year total cost25352
Implementation effort20345
Support fit10434
Data portability10433
Weighted total / 10010072.5076.2565.00

Test how easily the recommendation changes

Now move 10 weight points from cost to workflow fit: the weights become 45, 15, 20, 10, and 10. Keep all observed scores unchanged. A rises to 77.50, B falls to 73.75, and C reaches 70.00. The winner changes because A’s workflow advantage matters more and B’s cost advantage matters less. The team needs to settle that tradeoff before signing.

Also test disputed scores. If B’s workflow score falls from 4 to 3 after a realistic trial, its baseline total loses 8.75 points and becomes 67.50. This identifies a valuable next test: resolve the workflow evidence before spending another week comparing minor features. Preserve the baseline and sensitivity versions so the decision record explains both.

NASA’s decision-analysis guidance stresses documenting assumptions, uncertainty, and the limits of the evaluation method when recommending among alternatives. For a vendor choice, this means keeping the unresolved questions and the conditions that could reverse the recommendation beside the final score.[2]

Turn the comparison into a purchase decision

Finish with a short recommendation: preferred vendor, decisive evidence, accepted compromises, unresolved conditions, and the person who can approve the commitment. If rankings are close, name the additional test that would clarify the choice. If evidence remains missing, withhold a final total rather than rescaling completed criteria to 100; rescaling makes an incomplete proposal appear more complete than it is.

Carry the accepted requirements into implementation. Assign an owner to confirm the import, workflow, support contact, and export behavior after setup. Compare the actual first invoice with the evaluation assumptions. The scorecard becomes useful again when a renewal arrives because the original expectations remain visible.

Kona can help organize approved proposals into an evidence table and draft the comparison. Ask it to preserve blanks, cite the provided document and section for every score suggestion, and show the calculation. Review the source evidence yourself before accepting a score. The downloadable template below can also be completed without AI.

Yours to use

Vendor evaluation worksheet

A Markdown worksheet with mandatory checks, evidence fields, scoring rules, the complete numerical example, and a decision record.

Markdown · Opens in any text editor · No signup required

Download template

# Vendor evaluation scorecard

## Decision brief
- Purchase / workflow:
- Decision owner:
- Users and usage assumptions:
- Budget ceiling and evaluation period:
- Evaluation date:
- Current-process alternative:

## Mandatory checks — complete before ranking
| Requirement | Evidence needed | Vendor A | Vendor B | Vendor C | Evaluator / deadline |
|---|---|---|---|---|---|
| [Observable requirement] | [Test or document] | Pending | Pending | Pending | [Owner / date] |

Use Pass, Fail, or Pending. Do not approve with unresolved gates.

## Evidence register
| Vendor | Criterion | Observation | Document / section / test | Date | Open question | Owner |
|---|---|---|---|---|---|---|
| [Name] | [Criterion] | [Observed result] | [Reference] | [Date] | [Question] | [Owner] |

## Scoring policy
Use 1–5 only when evidence is sufficient. Leave unknown scores blank.
Write criterion-specific definitions for 1, 3, and 5 before evaluating.
Weights must total 100. Weighted points = weight * (score - 1) / 4.
Do not produce a final total while any weighted score is missing.

| Criterion | Weight | Score 1 anchor | Score 3 anchor | Score 5 anchor |
|---|---|---|---|---|
| Workflow fit | 35 | [Define] | [Define] | [Define] |
| First-year total cost | 25 | Above $21,000 | Above $15,000 to $18,000 | $12,000 or below |
| Implementation effort | 20 | [Define] | [Define] | [Define] |
| Support fit | 10 | [Define] | [Define] | [Define] |
| Data portability | 10 | [Define] | [Define] | [Define] |

Example cost score 4: above $12,000 to $15,000; score 2: above $18,000 to $21,000.
All costs and bands below are fictional. Replace them with approved purchase assumptions.

## Worked example — fictional vendors; all gates passed
| Criterion | Weight | Vendor A | Vendor B | Vendor C |
|---|---|---|---|---|
| Workflow fit | 35 | 5 | 4 | 4 |
| First-year total cost | 25 | 3 | 5 | 2 |
| Implementation effort | 20 | 3 | 4 | 5 |
| Support fit | 10 | 4 | 3 | 4 |
| Data portability | 10 | 4 | 3 | 3 |
| Weighted total / 100 | 100 | 72.50 | 76.25 | 65.00 |

First-year costs: A $16,800; B $11,400; C $19,200.
Vendor B total: 35*3/4 + 25*4/4 + 20*3/4 + 10*2/4 + 10*2/4 = 76.25.
Sensitivity weights: 45, 15, 20, 10, 10. Result: A 77.50; B 73.75; C 70.00.
Disputed score test: B workflow fit falls to 3 under baseline weights; B becomes 67.50.

## Decision record
- Preferred vendor and decisive evidence:
- Accepted tradeoffs:
- Unresolved conditions / owner / deadline:
- Sensitivity finding:
- Approval owner and date:
- Implementation acceptance checks:
- Renewal review date:

Sources

Sources and further reading

References are linked next to the claims they support. Worked examples illustrate the method; they are not Kona customer results or industry benchmarks.
  1. [1]

    Evaluation Tools

    Scottish Government Procurement Journey

  2. [2]

Editorial method

This guide combines the linked sources with a worked example and a reusable template. Replace example inputs with your own evidence and check the result before making a business decision.

About Kona’s editorial standards

Put it into practice

Adapt this to your business in Workspace

Copy the brief below, open Workspace, and paste it with the inputs you want to use. Review the assumptions and calculations before relying on the output.

Help me evaluate vendors using only the proposals and test notes I provide. Start with the purchase brief and mandatory pass/fail/pending requirements. Propose distinct criteria with weights totaling 100 and criterion-specific 1–5 scoring anchors for my approval. Keep unknown evidence blank. Cite the source document and section beside each proposed score. Calculate points as weight × (score − 1) ÷ 4, withhold final totals for incomplete vendors, and test a reasonable change in weights. Draft a recommendation with accepted tradeoffs and unresolved conditions. Do not invent vendor capabilities, prices, certifications, or test results.

Start in Workspace

FAQ

Answers to keep your planning sprint moving

Quick explanations and definitions you can share with your team when reviewing the research.

How many criteria should a vendor scorecard have?
Use enough criteria to represent the actual tradeoffs without counting the same benefit repeatedly. The worked example uses five, plus separate mandatory requirements. Add a criterion only when it could reasonably change the decision.
What do I do when a vendor has not supplied evidence?
Mark the requirement pending or leave the score blank, record the missing evidence, and assign a follow-up owner and deadline. Do not invent a score or normalize an incomplete comparison into a final ranking.
Should the highest-scoring vendor always win?
A score supports the decision under explicit assumptions. Check mandatory requirements, sensitivity, unresolved evidence, and accepted tradeoffs before approval. Document the reason for any departure from the ranking.