graph-studio
Matching Strategies
Exact Match (Default)
Finds connections where property values between two classes are identical.
How it works:
Builds a pivot layer that extracts all datatype property values from all classes into a flat structure (
class,property,value).For each class pair, executes a SPARQL query that joins source and target property values on exact equality (
?srcValue = ?dstValue), grouped by source and target properties.Returns the count of matching values per property pair.
Supports both string and numeric datatypes (
xsd:string,xsd:int,xsd:float,xsd:double).
Best for: Datasets that share identifiers, codes, or categorical values with consistent formatting (e.g., country codes, product IDs, zip codes).
Fuzzy Match
Finds connections where property values are similar but not necessarily identical, using an edit-distance threshold.
How it works:
Builds the same pivot layer as Exact Match.
For each source/target property pair, invokes the PPJoin (Prefix-filtered Partition Join) text similarity algorithm via Graph Lakehouse's built-in
jedai/text_similarityservice.PPJoin efficiently filters candidate pairs using prefix-based pruning, then computes string similarity scores. Only pairs exceeding the configured edit distance threshold (default 0.80, i.e., 80% similarity) are considered matches.
Returns the count of matching source and target instances per property pair.
Operates on
xsd:stringvalues only.
Best for: Datasets where value formats differ slightly — e.g., company names ("Cambridge Semantics" vs. "Cambridge Semantics Inc."), addresses with minor formatting differences, or names with typos.
Key parameter:
| Parameter | UI Label | Default | Description |
|---|---|---|---|
score |
editDistance |
Similarity Score | Minimum similarity threshold (0–100 in UI, converted to 0.0–1.0). Higher values require closer matches. |
Regex Match
Finds connections by first extracting a portion of each property value using a regular expression, then matching on the extracted values.
How it works:
Accepts separate source regex and target regex patterns.
Standard mode: Builds a specialized pivot layer that applies
regex_extractduring materialization — source values are extracted with the source regex and stored as:srcVal, target values with the target regex as:dstVal. The matching query then joins on the extracted values.Low Memory mode: Builds a standard pivot layer (like Exact Match) and applies
regex_extractat query time during pair matching, comparing extracted values rather than raw values.The
regex_extractservice is invoked via the Graph Lakehouseutil:regex_extractUDF, which pulls the first capturing group from each value.
Best for: Datasets where the connection exists within a substring of the values — e.g., extracting area codes from phone numbers, extracting domains from email addresses, or normalizing identifiers with different prefix/suffix patterns.
Key parameters:
| Parameter | UI Label | Description |
|---|---|---|
sourceRegex |
Source Regex | Regex applied to source class property values. Uses the first capture group. |
targetRegex |
Target Regex | Regex applied to target class property values. Uses the first capture group. |
lowMemoryUsage |
Use Low Memory Queries | When enabled, applies regex at query time instead of materializing extracted values in the pivot layer. Slower but uses less memory. |
Source: https://docs.sw.siemens.com/documentation/external/PL20260212925461721/en-US/graph_studio/find-graphmart-match-strat.htm · retrieved 2026-08-23