TEXTSIMILARITY

Applies to: Calculated column Calculated table Measure Visual calculation

Returns a floating-point relevance score that indicates how strongly text in a column matches a search string.

Syntax

TEXTSIMILARITY(<within_text_column>, <search_text>[, <match_mode>])

Parameters

Term Definition
within_text_column A column that contains the text values to search. The column must have the String data type.
search_text A text string that defines the search condition.
match_mode (Optional) The matching behavior to use. Supported values are TEXTMATCHING, FUZZYMATCHING, and PHRASEMATCHING. The default is TEXTMATCHING.

Return value

A floating-point lexical relevance score. Within one evaluation, a higher score indicates a stronger match, and 0 indicates no match.

The score is an opaque ranking value, not a percentage, probability, or normalized similarity measure. Compare scores only within the same search context.

Scores aren't guaranteed to be comparable or stable across separate query executions, search strings, columns, filter contexts, data sources or storage modes, index contents or generations, or execution plans. These factors can change the scoring corpus or scoring implementation even when visible text values are unchanged.

Match modes

Mode Behavior
TEXTMATCHING Tokenizes search_text and scores text that matches any search token.
FUZZYMATCHING Scores tokens that are similar to search tokens within a small edit distance. This mode supports typo-tolerant matching.
PHRASEMATCHING Scores text when the analyzed tokens appear next to each other and in the specified order. Punctuation doesn't affect the token sequence.

Remarks

  • The search is case-insensitive.
  • The function applies language-specific stemming and stop-word removal based on the model culture. For example, stemming allows a search for run to match running, runs, and runner. Stop-word removal ignores common filler words like the, and, of, and in.
  • Per-column collation doesn't control the analyzer language. The model culture controls it. In Power BI Desktop, select File > Options and settings > Options > Current file > Regional settings > Model language to set the model language. For more information, see FORMAT.
  • The supported model cultures are:
    • English (en)
    • French (fr)
    • German (de)
    • Spanish (es)
    • Portuguese (pt)
    • Russian (ru)
    • Arabic (ar)
    • Turkish (tr)
    • Danish (da)
    • Dutch (nl)
    • Finnish (fi)
    • Greek (el)
    • Hungarian (hu)
    • Italian (it)
    • Norwegian (no, nb, and nn)
    • Romanian (ro)
    • Swedish (sv)
    • Tamil (ta)
  • For unsupported cultures, TEXTSIMILARITY uses a generic lowercase tokenizer without stemming or stop-word removal.
  • Accent folding isn't supported. For example, cafĂ© doesn't match cafe.
  • The function requires a persisted full-text index on within_text_column. It's supported for Import, Dual, and Direct Lake storage modes. Other storage modes return an error.
  • The function is evaluated row by row. Before using it in a calculated column over a very large table, test the performance with a representative workload.
  • Use TEXTCONTAINS when you need a Boolean match result instead of a relevance score.

Examples

Return the most relevant rows

The following query returns the five reviews with the highest relevance scores for great refrigerator:

EVALUATE
TOPN (
    5,
    FactReviews,
    TEXTSIMILARITY ( FactReviews[Text], "great refrigerator" ), DESC
)
Example text Score
This is a great refrigerator, keeps everything cold and quiet. 3.42
Fantastic fridge! Great cooling performance and very spacious. 2.98
Our old refrigerator broke and this new one works great. 2.54
Refrigerator is decent, not great but gets the job done. 1.72
The refrigerator section is large, freezer is okay. 1.10

Filter and rank rows

The following query restricts the result to reviews that match refrigerator, and then ranks them by their relevance to great refrigerator:

EVALUATE
TOPN (
    5,
    FILTER (
        FactReviews,
        TEXTCONTAINS ( FactReviews[Text], "refrigerator" )
    ),
    TEXTSIMILARITY ( FactReviews[Text], "great refrigerator" ), DESC
)

TEXTCONTAINS and TEXTSIMILARITY execute independently. Filtering rows with TEXTCONTAINS doesn't change the index scope used to calculate TEXTSIMILARITY scores.

Example text Score
This is a great refrigerator - quiet, spacious, and keeps food fresh. 3.42
Fantastic fridge. Cooling performance is excellent and build quality is great. 3.01
Our old refrigerator died, but this new one works great and runs silently. 2.67
The refrigerator section is huge. Would be great if the freezer were larger. 1.94
Solid refrigerator. Not the best, but overall a great value for the price. 1.65

Display relevance scores

The following query adds the relevance score for each product description to the result:

EVALUATE
ADDCOLUMNS (
    Products,
    "Relevance score", TEXTSIMILARITY ( Products[Description], "wireless bluetooth" )
)
Example text Score
Wireless headphones with Bluetooth 2.35
Wired headphones 0