Skip to content

Privacy-Safe Data Fusion in Snowflake for Marketers

Toni Lee Cheib
Toni Lee Cheib

Key Takeaways

  • Privacy-safe data fusion connects fragmented datasets without requiring personal identifiers, enabling richer audience insights while maintaining compliance.
  • Probabilistic matching estimates which records belong together based on shared characteristics when no exact identifier exists between datasets.
  • Snowflake's environment enables marketers to run data fusion workflows in governed, secure spaces without moving sensitive data between organisations.
  • RSMB Fusion’s Snowflake Native App delivers privacy-safe data integration using probabilistic matching, helping marketers create measurement-ready datasets within their own Snowflake environments.
  • The real value of data fusion lies in answering business questions that no single dataset can address alone.

 

The value lies in connecting the data

Many organisations now hold substantial customer, campaign and research data in Snowflake. The difficulty is that these sources often remain separate, making it hard to answer questions that depend on seeing them together.

A CRM can highlight high-value customers, while campaign data shows who has been exposed to which media and how they responded. Research datasets add context on attitudes, interests and media habits. Each source is useful on its own, but the most important questions often require them to be considered together:

    • Which channels are most likely to reach a particular customer group?
    • How do attitudes relate to purchase behaviour?
    • What can first-party data tell us when it is enriched with trusted research sources?

Where the datasets do not share reliable identifiers, probabilistic data fusion can provide a way to connect them. The goal is to answer questions that a single dataset cannot support on its own.

 

Why marketing data remains fragmented

Bringing data into Snowflake does not automatically make it connected. Customer records may sit alongside campaign exposure, transaction data, website behaviour and third-party research, but each source is often structured differently and uses its own way of identifying individuals.

A CRM might rely on customer numbers or email addresses, while an advertising platform may use device or platform identifiers. Research datasets are often anonymised and may not include personal identifiers at all.

Even where two sources contain the same type of identifier, the overlap may be incomplete. People use different email addresses across services, records become outdated, and privacy or governance rules may prevent identifiers from being shared.

The data may therefore be held in the same environment but still cannot be analysed as a connected whole.

 

How probabilistic data fusion helps

Where reliable common identifiers exist, deterministic matching will usually be the most appropriate approach. Under deterministic matching, records are connected when they share an exact value, such as a customer number or email address.

Probabilistic fusion is useful where those identifiers are missing, inconsistent or unavailable. Instead of requiring an exact match, it uses characteristics that appear in both datasets to identify statistically similar records.

These shared characteristics, often known as fusion hooks, might include age, location, household composition, income, purchasing behaviour or media consumption. The method assesses the available information across several variables rather than relying on a single identifier.

Probabilistic fusion does not claim that two records definitely relate to the same individual. It identifies credible statistical matches that allow information to be transferred between datasets in a controlled way. This provides a practical route to combining sources that could not otherwise be linked.

 

What questions can connected data answer?

Fusion can support audience enrichment, segmentation, cross-media planning and measurement. Its value lies in the questions it allows an organisation to answer.

A business may want to understand which media channels are most likely to reach its highest-value customers. First-party data can show who those customers are and what they buy, but often provides limited insight into their wider media behaviour.

Fusing this data with a trusted planning or research source can add information on television viewing, online activity, radio listening, out-of-home exposure and other media habits.

This could help an organisation understand:

    • which media environments are most likely to reach different customer groups
    • how attitudes and interests vary by customer value or product holding
    • what distinguishes customers who respond to a campaign
    • where there are gaps or duplication in reach across channels

Fusion does not answer every question on its own. Connecting exposure and outcome data, for example, does not prove that advertising caused a sale. Establishing causality or incrementality requires an appropriate experimental or analytical design alongside the connected data.

 

A practical example

Consider a financial services provider that holds detailed customer, product and transaction data in Snowflake but has little understanding of its customers’ wider media behaviour.

An industry planning dataset such as IPA TouchPoints or Barb contains rich information on media use. However, it cannot be linked directly to the customer data because the two sources do not share personal identifiers.

Probabilistic fusion could use characteristics present in both datasets, such as age, region and household income, to identify statistically similar records.

The resulting dataset would not claim that a particular customer watches a particular television programme. It could, however, provide a robust view of the media behaviours associated with different customer groups.

This could help the organisation plan media around high-value segments, enrich its audience profiles and make better-informed investment decisions.

The quality of the result would depend on the source data, the strength of the shared variables and the validation applied to the fusion.

 

When is probabilistic fusion appropriate?

Probabilistic fusion is most likely to be useful where:

    • the question depends on more than one dataset
    • exact identifiers are unavailable or provide insufficient coverage
    • the sources contain meaningful variables in common
    • a statistical connection, rather than a confirmed individual-level link, is appropriate for the intended use

The shared variables need to do more than simply appear in both sources. They should help distinguish between different types of people or behaviour. Age and region may contribute, but are unlikely to be sufficient on their own.

Variables can be weighted according to their relevance, while critical characteristics can be required to match exactly where the analysis requires it.

Fusion also introduces uncertainty. This may be acceptable for planning, segmentation and some forms of measurement. It may not be suitable where a decision depends on a confirmed link to a named individual.

 

What makes a fusion credible?

The number of records connected is not a reliable measure of quality. A method could create a match for every record and still produce a poor analytical dataset. The more important question is whether the result preserves the patterns and relationships needed for its intended use.

Three factors are particularly important.

1. The quality of the source data

Fusion cannot repair fundamental weaknesses in the underlying sources.

Poor coverage, inconsistent definitions, high levels of missing data and unrepresentative samples can all affect the result. The source datasets should therefore be assessed before any matching takes place.

2. The choice of fusion hooks

The variables used for matching should be relevant to the information being transferred.

A fusion intended to add media behaviour to customer records, for example, should use characteristics that help predict differences in media consumption.

Variables should also be measured consistently across the sources. Two fields with similar names may not represent the same thing.

3. The validation of the output

A fusion should be tested to establish whether it has produced a credible result.

Validation may include checking whether:

    • important distributions from the source data have been preserved
    • relationships between variables remain plausible
    • matched records are similar on measures not used in the matching
    • results remain stable when assumptions or matching variables change
    • the fused dataset performs adequately for its intended use

It is relatively straightforward to produce a connected dataset. Demonstrating that it deserves confidence requires more work.

 

Constrained and unconstrained fusion

Different projects require different forms of fusion.

Unconstrained fusion seeks to find the best available statistical match for each record based on the selected variables. This can work well where the individual matches also produce an acceptable overall result.

However, the best individual matches may still distort an important pattern or distribution. For example, a fusion might produce reasonable record-level matches but misrepresent the relationship between age, customer value and media behaviour.

Constrained fusion addresses this by controlling the aggregate structure of the fused dataset, helping it remain aligned with known totals or relationships.

This is particularly important in audience and media analysis, where users need confidence that the overall dataset remains representative and that key patterns in the source data have not been lost.

 

Why run fusion within Snowflake?

For organisations already using Snowflake, carrying out the process within the same environment avoids unnecessary data movement and allows existing governance, permissions and security controls to remain in place.

It also makes it easier to work with first-party data alongside research, media or partner datasets made available through controlled sharing arrangements.

Snowflake provides the environment in which the data can be stored, governed and analysed. It does not, on its own, solve the problem of connecting records where reliable identifiers are missing. That requires an appropriate matching method, supported by careful validation.

RSMB Fusion runs as a Snowflake Native App within the customer’s Snowflake environment. The data remains under the customer’s control while the matching and validation take place within the platform.

This can reduce some of the security, legal and governance work associated with transferring data to an external processing environment. The precise requirements will still depend on the datasets and their intended use.

Probabilistic fusion can also reduce the need to exchange personal identifiers such as email addresses, customer numbers or device IDs.

This does not remove wider privacy and compliance obligations. Organisations must still consider the purpose and lawful basis for processing, access controls, the sensitivity of the data and whether individuals could be identified from the resulting combination of attributes.

 

A sound approach to data fusion

A fusion project should begin with a clear business question and an assessment of whether the available data can support it.

The source datasets and potential fusion hooks can then be reviewed before the matching method, any required constraints and the validation criteria are agreed.

Once the fusion has been run, the resulting dataset should be tested against the original sources and its intended use. The technical matching may be relatively quick, but data preparation, methodological design and validation usually require more judgement.

Any provider should be able to explain how the matching works, which variables drive the result, where uncertainty enters the process and how the output is validated. The test is not how many records have been connected, but whether the resulting dataset provides a reliable basis for analysis and decision-making.


 

How RSMB Fusion can help

RSMB has delivered high quality audience measurement, data integration and statistical modelling for more than 35 years.

That background has shaped RSMB Fusion, particularly its focus on methodological transparency and validation.

RSMB Fusion supports both constrained and unconstrained approaches. It includes validation tools that allow users to assess the quality of the fused dataset and understand whether it is suitable for its intended purpose.

As a Snowflake Native App, it allows organisations to carry out the fusion within their existing environment rather than transferring the data elsewhere for processing.

The benefit is the ability to answer useful questions across customer, campaign and research data while maintaining appropriate control over privacy, governance and statistical quality.

Where exact identifiers are unavailable or provide insufficient coverage, probabilistic fusion may offer a practical way to make more of the data already held in Snowflake.

RSMB helps organisations assess, design and deliver data fusion projects within Snowflake. To discuss whether fusion could help answer a marketing or insights question, get in touch.

Share this post