GraphKnowledge

graph-studio

Matching Strategies

Exact Match (Default)

Finds connections where property values between two classes are identical.

How it works:

  • Builds a pivot layer that extracts all datatype property values from all classes into a flat structure (class, property, value).

  • For each class pair, executes a SPARQL query that joins source and target property values on exact equality (?srcValue = ?dstValue), grouped by source and target properties.

  • Returns the count of matching values per property pair.

  • Supports both string and numeric datatypes (xsd:string, xsd:int, xsd:float, xsd:double).

Best for: Datasets that share identifiers, codes, or categorical values with consistent formatting (e.g., country codes, product IDs, zip codes).

Fuzzy Match

Finds connections where property values are similar but not necessarily identical, using an edit-distance threshold.

How it works:

  • Builds the same pivot layer as Exact Match.

  • For each source/target property pair, invokes the PPJoin (Prefix-filtered Partition Join) text similarity algorithm via Graph Lakehouse's built-in jedai/text_similarity service.

  • PPJoin efficiently filters candidate pairs using prefix-based pruning, then computes string similarity scores. Only pairs exceeding the configured edit distance threshold (default 0.80, i.e., 80% similarity) are considered matches.

  • Returns the count of matching source and target instances per property pair.

  • Operates on xsd:string values only.

Best for: Datasets where value formats differ slightly — e.g., company names ("Cambridge Semantics" vs. "Cambridge Semantics Inc."), addresses with minor formatting differences, or names with typos.

Key parameter:

Parameter UI Label Default Description
score editDistance Similarity Score Minimum similarity threshold (0–100 in UI, converted to 0.0–1.0). Higher values require closer matches.

Regex Match

Finds connections by first extracting a portion of each property value using a regular expression, then matching on the extracted values.

How it works:

  • Accepts separate source regex and target regex patterns.

  • Standard mode: Builds a specialized pivot layer that applies regex_extract during materialization — source values are extracted with the source regex and stored as :srcVal, target values with the target regex as :dstVal. The matching query then joins on the extracted values.

  • Low Memory mode: Builds a standard pivot layer (like Exact Match) and applies regex_extract at query time during pair matching, comparing extracted values rather than raw values.

  • The regex_extract service is invoked via the Graph Lakehouse util:regex_extract UDF, which pulls the first capturing group from each value.

Best for: Datasets where the connection exists within a substring of the values — e.g., extracting area codes from phone numbers, extracting domains from email addresses, or normalizing identifiers with different prefix/suffix patterns.

Key parameters:

Parameter UI Label Description
sourceRegex Source Regex Regex applied to source class property values. Uses the first capture group.
targetRegex Target Regex Regex applied to target class property values. Uses the first capture group.
lowMemoryUsage Use Low Memory Queries When enabled, applies regex at query time instead of materializing extracted values in the pivot layer. Slower but uses less memory.

Source: https://docs.sw.siemens.com/documentation/external/PL20260212925461721/en-US/graph_studio/find-graphmart-match-strat.htm · retrieved 2026-08-23