Skip to content

How to decide on weighting a match score you can defend

MatchmakingUpdated 2026-08-188 min read

In short

Weighting a match score means choosing how much each sub-score contributes to the total, and the choice should be tested against held out proposals from a prior edition rather than argued in a meeting. Compare candidate weight sets on precision at five proposals per buyer, then publish the winning weights.

The argument always happens in the same meeting. The sales director wants product fit to count for more, because exhibitors complain about product mismatches and never complain about industry mismatches. The data lead wants intent to count for more, because intent is the only term that moved last year. Somebody proposes splitting the difference, and the weights end up as whatever the loudest person said, written nowhere.

Weighting a match score is a policy choice, and policy choices are allowed. What is not allowed is making one and then describing the output as if a model produced it. The fix is small: hold out a set of proposals whose outcome you already know, score them under each candidate weight set, and let the numbers narrow the argument before opinion finishes it.

Why are the weights a policy choice?

Because there is no external fact that fixes them. The five sub-scores are on the same scale by construction, and how much each one should count depends on what your show is for.

A hosted buyer programme at a sourcing show has one job, which is putting people who can sign a purchase order in front of suppliers who can fill it. Authority fit matters there in a way it does not at a technology conference where the value is the conversation. A first edition has no behavioural history at all, so weighting behaviour at 0.13 means weighting a mostly empty column at 0.13.

Writing the weights down as policy has a second benefit. It gives the sales director something to disagree with in a form that can be tested, which is better than a disagreement about vibes that resurfaces every edition.

Building the held out set

You need proposals with outcomes attached. Last edition's meeting programme has them, and most organisers have never assembled them into one file.

One row per proposal shown to a buyer. The five sub-scores as they were at the time of proposal, the identity of both parties, the reason code that led the proposal if you keep one, which is I5's subject, and the outcome. Outcome needs care, because there are three of them and they mean different things. Accepted and held is the one you are predicting. Accepted and no-show is a behaviour problem sitting inside a match that worked. Declined is the negative case. Never shown is not an outcome at all, and including unshown pairs as negatives will teach your weights that your old ranking was correct.

Recomputing the sub-scores as they stood at proposal time is the part that gets skipped, and skipping it leaks the future into the test. A buyer who accepted three meetings has a higher intent score afterwards than they did at the moment the fourth proposal was made. Score from a snapshot, or accept that your test is flattering you.

Take a realistic file. 214 hosted buyers, five proposals each, 1,070 proposals, 468 of them accepted. The base rate is 468 divided by 1,070, which is 0.437.

What to measure, and why precision at five

Precision at five is the share of the top five proposals per buyer that get accepted, averaged across buyers. It is the right measure here because it matches what the system does. Your concierge team is not ranking the whole exhibitor list for each buyer. They are putting five names in front of somebody with a twelve meeting quota, and the only question is how many of those five survive.

Recall is the wrong measure for the same reason. A buyer with a twelve meeting quota cannot take 40 good matches, so finding all of them earns nothing. Precision at the depth you actually show is the number that changes a decision.

Two other measures are worth carrying alongside. Coverage, which is the share of exhibitors appearing in anybody's top five, because a weight set that sends every buyer to the same 30 stands scores well and ruins the floor, a failure that belongs to I39. And the share of buyers who filled their quota, because a weight set that produces very few proposals above threshold has optimised itself into a scheduling problem.

Running the grid and reading it honestly

Five weights summing to one, moved in steps of 0.05, gives 10,626 candidate weight vectors. Scoring 1,070 proposals under each is a few seconds of compute, which is what makes this tempting and dangerous.

Run three of them by hand first, because the hand version is the one you can explain. Flat weights of 0.2 each. A product-heavy set of 0.45 product, 0.15 industry, 0.20 intent, 0.10 authority and 0.10 behaviour. And the config you are running today, 0.30, 0.20, 0.22, 0.15 and 0.13.

Suppose flat weights return precision at five of 0.44, the product-heavy set returns 0.47, and the current config returns 0.46. The product-heavy set looks like the winner and the gap is three points, which on 1,070 proposals sounds like 32 extra accepted meetings.

Now compute the error band. The standard error on a proportion near 0.45 across 1,070 observations is the square root of 0.45 times 0.55 divided by 1,070, which is the square root of 0.000231, which is 0.0152. So 1.5 points, and the three point gap is roughly two standard errors. Except the proposals are not independent: they arrive in groups of five from 214 buyers, and buyers differ from each other far more than proposals within a buyer do. A design effect of two on that clustering pushes the standard error to about 2.2 points, and the three point gap falls to 1.4 standard errors, which is not a result.

The 10,626 point grid makes this worse. The best vector out of 10,626 evaluated on one held out set will beat the flat baseline by several points on that set and much less on the next one, because you have selected on noise. Split the file into two halves by buyer, fit on one, report on the other, and quote the second number only.

What does the flat 0.2 baseline actually cost you?

Less than the meeting assumes, and this is the most useful thing in the literature for anyone running this test.

Dawes argued in American Psychologist in 1979 that improper linear models, including ones with equal weights, predict about as well as optimally weighted ones across a long list of applied problems, and that knowing which variables belong in the model and in which direction matters more than the weights on them. Einhorn and Hogarth had set out the mechanism in Organizational Behavior and Human Performance in 1975: the correlation between a unit weighted composite and a regression weighted one rises as the predictors correlate with each other and falls as the number of predictors grows.

Both conditions hold in a match score. There are only five terms, and product fit and industry fit are heavily correlated because they are both read off the same taxonomy, while intent and behaviour correlate through the simple fact that engaged buyers do more of everything. A five term score built from correlated terms is close to the worst case for elaborate weighting and close to the best case for a flat split.

The practical consequence is a rule for the meeting. Depart from equal weights when the held out difference clears your error band, and keep the departure small. Moving product fit from 0.20 to 0.30 is a defensible policy statement. Moving it to 0.45 on the strength of a three point difference is an unforced error.

Writing the weights down where somebody can object

Weights belong in a versioned config file with a date, an owner and a one line reason for each value, next to the held out precision figure they were chosen on.

That file is also what makes the next edition tractable. When precision at five drops from 0.46 to 0.39, the first question is whether the weights changed, and without a versioned record that question takes a week. The second question is whether the sub-scores changed underneath the weights, which happens when somebody rebuilds the taxonomy and product fit quietly starts meaning something else.

Publishing the weights to exhibitors is a separate decision with its own consequences, and the argument for coarse bands and a held back behaviour term sits in what you show the exhibitor, which is I6.

Where the grid search stops

Everything above optimises for accepted proposals, which is a proxy. The thing you want is meetings that produce business, and acceptance is what you can measure four weeks after the show without chasing anybody.

The gap between the two is real and it points in a known direction. Buyers accept meetings with names they recognise, so a weight set optimised on acceptance will drift towards large well known exhibitors, and your smallest exhibitors will lose proposals for a reason that has nothing to do with fit. If you have outcome data from post meeting surveys, weight against that instead and accept the smaller sample.

The second limit is that a single weight set is being asked to serve buyers who want different things. A returning buyer with a full profile and a first-time buyer with three fields completed are being ranked by the same formula, and the terms available to them differ. Segmenting the weights by profile completeness is defensible, doubles the number of things to test, and is a reasonable second year project.

Start this week by building the held out file. One row per proposal from last edition, five sub-score columns as they stood on the day, and the outcome. Compute precision at five under flat weights and under whatever you are running now. If the two numbers are within your error band, you have learned that the weights are not where your improvement is, and the next place to look is either what the sub-scores contain in I1 or whether the scores mean what they say in I3. Both are better places to spend a fortnight than a grid search, and both feed the same matchmaking reporting.

Questions people ask about weighting a match score

How do you choose weights for a matchmaking score?
Hold out the proposals from a prior edition with known accept and decline outcomes, then score them under each candidate weight set and compare precision at five proposals per buyer. Pick the weights that win by more than the sampling error of your own file, and treat anything smaller as a tie.
Are equal weights good enough for a match score?
Often, yes. Dawes argued in American Psychologist in 1979 that equal weighting performs close to optimised weighting across many prediction problems, and the effect is stronger when the predictors correlate with each other. Product fit and industry fit correlate heavily, so a flat split is a serious baseline rather than a placeholder.
How much data do you need to test match score weights?
Enough that the difference you care about is larger than the standard error. With around 1,000 proposals the standard error on a precision figure near 0.45 is about 1.5 points before any clustering adjustment, so a two point improvement is not evidence. Proposals cluster within buyers, which widens that band further.

Related reading

All matchmaking articles