Prospect lists look reassuringly precise. Ten thousand rows feel like ten thousand prospects, each neatly captured with a company name, website, contact and category. Lists form the bedrock of almost any launch, although grown-up organisers tend to hide them as much as possible inside sales and marketing systems. Most organisers are realistically drowning in lists, with technology often the only way to achieve some kind of order. Lists are created, bought, combined, used and abused throughout the show cycle.

“The number of rows can then become a kind of vanity measure for how successful you are”

The number of rows can then become a kind of vanity measure for how successful you are. Count the rows and you know how much data you have. Except you don’t, because none of those rows is actually a contact, a customer or a company. It is an attempt to represent one: a small, imperfect model of something that exists outside the list. The reassuring certainty of the row disguises a much messier reality; a list is really just a data caterpillar, hiding a business butterfly within.

If the representation is poor, every campaign, analysis or tool built on top of it inherits the problem. Thinking purely in lists and rows is therefore a poor way of understanding data. Imagine a fictional potential exhibitor called Northstar Mobility Ltd. We might have a row containing its name, an address in Manchester, a website, a product category, an employee count of 620 and two appearances at a fictional show called FutureRail. Maybe, if we’re lucky, we also have some identified contacts and their email addresses.

This data represents attributes someone has recorded about Northstar. Some are measured, some calculated and some estimated. The list does not contain Northstar itself; it contains claims about Northstar. Before those claims become useful, however, we have to answer an apparently simple question: which Northstar? Another list may also contain records for Northstar Mobility, North Star Mobility Ltd and Northstar Group plc.

A process called “entity resolution”

Are they three businesses, two businesses and a spelling mistake, or three names for the same thing? We need to work out which records belong together and which are genuinely distinct. This is a process called “entity resolution”. Matching names, websites, addresses or company numbers can help, as can someone who knows the market. Yet there is not always one correct answer, because the boundary depends on what we mean by “company”.

Sales may treat a group as one account, while marketing may treat its trading brands separately. Finance needs the legal entity signing the contract, while attendees may recognise the brand rather than the company behind it. They are all connected, but they are not interchangeable. Merge too little and one prospect appears five times; merge too much and a parent, subsidiary and brand collapse into a single fictional blob. The supposedly simple act of “cleaning” the data is already becoming something rather more complicated.

“What looks like hygiene is only an interpretation”

Cleaning data is not merely tidying punctuation or standardising differences in wording. Every merge quietly makes a judgement about what exists, and every decision to keep two records apart does the same. Those choices may be hidden inside a formula or made manually, but they shape every result that follows. Many are destructive because, in forcing order onto data, we often turn a claim into a fact and discard the uncertainty around it. What looks like hygiene is only an interpretation.

From lists to networks

Once we know what the things are, however, we can start to connect them. “Northstar has 620 employees” is an attribute, while “Northstar exhibited at FutureRail 2026” is a relationship. So are “Aisha Khan works at Northstar Mobility” and “Northstar Mobility is part of Northstar Group”. This is where lists start to creak, because most lists are mixtures of attributes and relationships: data about things alongside data about the links between them.

Recognising this is a fundamental change in how we think about data. Viewed another way, the same information becomes a network, or what data nerds might call a “graph”: things joined to other things. That allows us to ask better questions. Which companies have exhibited at two competing shows but never at ours? Which contacts have moved between companies in the sector, and which corporate groups are present through several apparently separate brands? These are questions that help us understand markets rather than simply count records.

The value therefore sits not only in the things themselves, but in the connections between them. At this stage, we probably also realise that our original list contains only a partial set of claims about all the potentially relevant relationships. We are almost certainly missing some. Luckily, organisers rarely have just one list; they are more likely drowning in new data, old data, bought data, registration data, CRM data, public data and exhibitor data. Each offers a different perspective and, quite often, a competing claim.

When and where it came from

The “when and where” of a piece of data now become critical. Northstar may have had 620 employees in 2022 and 900 in 2026, and both figures can be correct. Aisha may have worked there until March and joined another business in April. If we keep only the latest value, we lose the change; if we ignore time altogether, yesterday’s truth eventually starts to look like today’s error. Where data comes from matters just as much as the value itself, because some sources provide stronger evidence than others.

“What you really have are 10,000 attempted representations”

So, when someone proudly announces that they have a 10,000-row prospect list, the list is telling a comforting story: here are 10,000 identifiable companies waiting to be worked. The reality is messier. What you really have are 10,000 attempted representations containing identities, attributes, boundaries, relationships, similarities and states through time. At every stage, someone has likely made choices that either sharpen the map or blur it. The real test is whether the model preserves the distinctions our decisions depend upon and makes its limitations visible.

If you think about data purely in rows and records, you may still have a perfectly serviceable mailing list. But you do not yet have a model of your market. Only when you understand the competing claims, the relationships between them and how they change through time can you begin to see the hidden patterns in your market data. That is where organiser data becomes much more interesting, because years of exhibitors, visitors, job moves, brands, ownership structures and show participation start to describe not simply who is in the market, but how the market itself works.

“The next time you open a prospect list, don’t simply ask how many rows it contains”

My advice is: the next time you open a prospect list, don’t simply ask how many rows it contains. Ask what claims those rows are making, what uncertainty they hide, how they connect with everything else you know, and whether those new claims strengthen or distort your model of the market. Your prospect list may not exactly be lying to you, but if you treat every row as the truth, you may end up lying to yourself.