Skip to content

Cold start matchmaking at a first edition with no behavioural history

MatchmakingUpdated 2026-08-187 min read

In short

Cold start matchmaking means scoring proposals at a launch edition where no accepted meetings exist. Hold the intent and behaviour terms at a neutral value, renormalise the remaining weights across product, industry and authority so they sum to one, and instrument the edition so the second one has outcome data.

A launch edition in a new vertical had 340 exhibitor contracts signed and 2,900 buyers registered six weeks out. The matchmaking platform had been bought on a demo where a model learned from three years of meeting outcomes. There were no meeting outcomes. There was one registration file and a spreadsheet of exhibitor product categories that the sales team had filled in while chasing contracts.

Cold start matchmaking at a first edition is the normal state for launches, for shows that changed platform, and for any organiser whose previous system never wrote an outcome back to a row. The score still has to run, buyers still expect proposals, and the answer is to be explicit about which terms are dead and to stop pretending otherwise.

What is actually missing at a launch edition?

Less than the panic suggests. Two of the five terms are missing, and three are intact.

Product fit comes from the exhibitor's declared categories against the buyer's declared interests, and both exist. Industry fit comes from the buyer's sector and the exhibitor's target sectors, which are in the contract and the registration form. Authority fit comes from the declared purchasing role, the employer size band and the hosted buyer screen, none of which need history.

Intent fit and behaviour fit are the casualties. Intent needs saves, views and requests, and at six weeks out a launch edition has a handful. Behaviour needs a prior edition, and there is no prior edition. Together those two terms carry 0.35 of the standard weighting, and the question is what to do with a third of your score.

The wrong answer, which is the default in most implementations, is to score them as zero. Zero is a measurement, and it says this buyer showed no intent, which is a claim about a buyer who never had the chance.

Renormalise the weights instead of scoring zeros

Set the two dead terms to a neutral 0.5 and you avoid the false claim, but you inherit a subtler problem: a constant is added to every candidate, which squashes the range without changing the order.

Work it through with the standard weights of 0.30 product, 0.20 industry, 0.22 intent, 0.15 authority and 0.13 behaviour. Take a strong candidate with product fit 0.90, industry fit 0.80 and authority fit 0.60. With intent and behaviour held at 0.5, the total is 0.30 times 0.90, plus 0.20 times 0.80, plus 0.22 times 0.50, plus 0.15 times 0.60, plus 0.13 times 0.50. That is 0.270 plus 0.160 plus 0.110 plus 0.090 plus 0.065, which is 0.695.

Now a weak candidate with product fit 0.40, industry fit 0.30 and authority fit 0.30. The total is 0.120 plus 0.060 plus 0.110 plus 0.045 plus 0.065, which is 0.400. The gap between the two is 0.295, and 0.175 of both scores is a constant contributed by terms measuring nothing.

Renormalise instead. Drop the dead terms and rescale the live ones so they still sum to one: 0.30 plus 0.20 plus 0.15 is 0.65, so product becomes 0.30 divided by 0.65, which is 0.462, industry becomes 0.308 and authority becomes 0.231. The strong candidate now scores 0.462 times 0.90, plus 0.308 times 0.80, plus 0.231 times 0.60, which is 0.416 plus 0.246 plus 0.138, or 0.800. The weak candidate scores 0.185 plus 0.092 plus 0.069, which is 0.346.

The gap widens from 0.295 to 0.454, and every number in the output now refers to something that was measured. Any threshold you set, any cut-off for a concierge review queue, any number shown to an exhibitor, is comparing candidates on evidence.

The trap to avoid is leaving the renormalised weights in place once behaviour data exists. Write the renormalisation as a function of which terms are live for that event instance, so the weights revert automatically at edition two, and keep the original policy weights in the config as the source of truth. Deciding what those policy weights should be once outcomes exist is a separate exercise, and it belongs to testing weights against held out proposals in I2.

Attributes are all you have, so the taxonomy is the product

When the whole score rests on declared attributes, the quality of the category tree stops being a data hygiene question and becomes the main determinant of match quality.

The recommender literature reached this conclusion by a different route. Gantner, Drumond, Freudenthaler, Rendle and Schmidt-Thieme addressed the case of no known prior events at ICDM in 2010 by mapping user and item attributes into the latent feature space of a factorisation model, so that attributes carry the prediction when interactions are absent. The event version needs none of that machinery, since your attributes are already the categories buyers and exhibitors picked, but the principle transfers: with no interactions, attribute quality is the ceiling on everything.

Two practical consequences follow at a launch. Exhibitor categories collected during contracting are usually filled in by a salesperson in a hurry, so they need a review pass before they become the product fit term for the whole show. And the buyer side needs enough category detail to discriminate, which is a direct trade against form completion and is the onboarding question problem in I18.

How do you evaluate a score with nothing to test against?

Not by backtesting, which is unavailable, and not by asserting the score is fine, which is what usually happens.

Schein, Popescul, Ungar and Pennock treated this directly at SIGIR in 2002, in work on methods and metrics for cold-start recommendations. Their contribution that matters most here is the evaluation half: they set out testing methodologies for the cold-start case and introduced the CROC curve as a performance measure, arguing that a single accuracy figure conceals how a recommender behaves on items nobody has rated. A launch edition is the same situation with exhibitors in place of items.

The practical translation is to build the evaluation into the show itself.

  • Log the state at proposal time. The five sub-scores, the rank, the threshold, the reason code and the origin, written when the proposal is created. Reconstructing this afterwards from current state is impossible, and every organiser who skips it discovers that in month two.
  • Randomise a slice. Take 5 per cent of proposals and fill them by random draw from candidates above a minimum product fit. That slice is the only unbiased sample you will ever have of what acceptance looks like without the score, and 5 per cent of 4,000 proposals is 200 rows, which is enough to see a gross failure.
  • Split the show in half by time. Day one produces accepted and declined meetings by lunchtime. Those are outcomes. A first calibration on day one's data, applied to day two's proposals, is a real experiment that costs nothing.

The randomised slice is the part organisers resist, and the objection is fair: you are deliberately proposing some meetings that score lower. The cost is bounded and known, the alternative is a second edition where you still cannot say whether the score helped, and 200 rows is a small price for the first honest number anyone has about your matchmaking.

What to hand the concierge team while the score is thin

At a launch, human review is doing more of the work than the model, and the interface should admit it.

Sort the queue by product fit alone for the first week and let the team see the ranking they would have made by hand. Where they disagree, capture the reason in a controlled list. Those disagreements are a labelled dataset built by people who know the vertical, and by the end of the show you will have a few hundred of them, which is more supervision than the model has from any other source.

Set the review threshold low enough that the team sees the borderline cases. If the score puts 60 per cent of proposals above the automatic threshold at a first edition, the threshold is wrong, because nothing in the file justifies that much confidence.

Where this stops

Declared attributes are claims. A buyer ticking eleven of fourteen categories on a registration form has told you nothing, and at a first edition there is no behavioural evidence to contradict them. Category selection inflation is worst exactly where it hurts most, among buyers keen enough to be in a hosted programme.

The second limit is that a first edition has no idea what its own acceptance rate should be. A 34 per cent acceptance rate on proposals might be excellent for a new vertical and poor for a mature one, and there is no internal benchmark. Priors from a sister event in the same portfolio are the only defensible source of one, and borrowing them properly is what I19 covers, including how much to trust a category with twelve local observations.

Before your launch opens, write the proposal log schema and check that it captures the sub-scores at proposal time. That is one table, an afternoon of work, and it is the difference between a second edition that can test its weights and a second edition that starts again from zero, whatever else your matchmaking configuration does this year.

Questions people ask about cold start matchmaking first edition

How do you run matchmaking at a first edition with no data?
Score on declared attributes only. Product fit, industry fit and authority fit all come from registration answers and exhibitor contracts, which exist before the doors open. Set the intent and behaviour terms to a neutral value and renormalise the three remaining weights so the total still sums to one.
Why renormalise the weights instead of scoring zero for missing terms?
A constant term adds the same amount to every candidate, which compresses the range of scores without changing their order. Renormalising restores the spread. In one worked case the gap between a strong and a weak candidate widened from 0.295 to 0.454 once the two dead terms were removed and the rest were rescaled.
How do you evaluate a matchmaking score with no history?
You cannot backtest, so instrument the edition instead. Log every proposal with its sub-scores and rank at the moment of proposal, randomise a small slice of proposals to create an unbiased sample, and treat the first day of accepted meetings as the first evaluation set for the second day.

Related reading

All matchmaking articles