Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Applies to:
Calculated column
Calculated table
Measure
Visual calculation
Returns a floating-point relevance score that indicates how strongly text in a column matches a search string.
Syntax
TEXTSIMILARITY(<within_text_column>, <search_text>[, <match_mode>])
Parameters
| Term | Definition |
|---|---|
within_text_column |
A column that contains the text values to search. The column must have the String data type. |
search_text |
A text string that defines the search condition. |
match_mode |
(Optional) The matching behavior to use. Supported values are TEXTMATCHING, FUZZYMATCHING, and PHRASEMATCHING. The default is TEXTMATCHING. |
Return value
A floating-point lexical relevance score. Within one evaluation, a higher score indicates a stronger match, and 0 indicates no match.
The score is an opaque ranking value, not a percentage, probability, or normalized similarity measure. Compare scores only within the same search context.
Scores aren't guaranteed to be comparable or stable across separate query executions, search strings, columns, filter contexts, data sources or storage modes, index contents or generations, or execution plans. These factors can change the scoring corpus or scoring implementation even when visible text values are unchanged.
Match modes
| Mode | Behavior |
|---|---|
TEXTMATCHING |
Tokenizes search_text and scores text that matches any search token. |
FUZZYMATCHING |
Scores tokens that are similar to search tokens within a small edit distance. This mode supports typo-tolerant matching. |
PHRASEMATCHING |
Scores text when the analyzed tokens appear next to each other and in the specified order. Punctuation doesn't affect the token sequence. |
Remarks
- The search is case-insensitive.
- The function applies language-specific stemming and stop-word removal based on the model culture. For example, stemming allows a search for
runto matchrunning,runs, andrunner. Stop-word removal ignores common filler words likethe,and,of, andin. - Per-column collation doesn't control the analyzer language. The model culture controls it. In Power BI Desktop, select File > Options and settings > Options > Current file > Regional settings > Model language to set the model language. For more information, see FORMAT.
- The supported model cultures are:
- English (
en) - French (
fr) - German (
de) - Spanish (
es) - Portuguese (
pt) - Russian (
ru) - Arabic (
ar) - Turkish (
tr) - Danish (
da) - Dutch (
nl) - Finnish (
fi) - Greek (
el) - Hungarian (
hu) - Italian (
it) - Norwegian (
no,nb, andnn) - Romanian (
ro) - Swedish (
sv) - Tamil (
ta)
- English (
- For unsupported cultures,
TEXTSIMILARITYuses a generic lowercase tokenizer without stemming or stop-word removal. - Accent folding isn't supported. For example,
cafédoesn't matchcafe. - The function requires a persisted full-text index on
within_text_column. It's supported for Import, Dual, and Direct Lake storage modes. Other storage modes return an error. - The function is evaluated row by row. Before using it in a calculated column over a very large table, test the performance with a representative workload.
- Use TEXTCONTAINS when you need a Boolean match result instead of a relevance score.
Examples
Return the most relevant rows
The following query returns the five reviews with the highest relevance scores for great refrigerator:
EVALUATE
TOPN (
5,
FactReviews,
TEXTSIMILARITY ( FactReviews[Text], "great refrigerator" ), DESC
)
| Example text | Score |
|---|---|
This is a great refrigerator, keeps everything cold and quiet. |
3.42 |
Fantastic fridge! Great cooling performance and very spacious. |
2.98 |
Our old refrigerator broke and this new one works great. |
2.54 |
Refrigerator is decent, not great but gets the job done. |
1.72 |
The refrigerator section is large, freezer is okay. |
1.10 |
Filter and rank rows
The following query restricts the result to reviews that match refrigerator, and then ranks them by their relevance to great refrigerator:
EVALUATE
TOPN (
5,
FILTER (
FactReviews,
TEXTCONTAINS ( FactReviews[Text], "refrigerator" )
),
TEXTSIMILARITY ( FactReviews[Text], "great refrigerator" ), DESC
)
TEXTCONTAINS and TEXTSIMILARITY execute independently. Filtering rows with TEXTCONTAINS doesn't change the index scope used to calculate TEXTSIMILARITY scores.
| Example text | Score |
|---|---|
This is a great refrigerator - quiet, spacious, and keeps food fresh. |
3.42 |
Fantastic fridge. Cooling performance is excellent and build quality is great. |
3.01 |
Our old refrigerator died, but this new one works great and runs silently. |
2.67 |
The refrigerator section is huge. Would be great if the freezer were larger. |
1.94 |
Solid refrigerator. Not the best, but overall a great value for the price. |
1.65 |
Display relevance scores
The following query adds the relevance score for each product description to the result:
EVALUATE
ADDCOLUMNS (
Products,
"Relevance score", TEXTSIMILARITY ( Products[Description], "wireless bluetooth" )
)
| Example text | Score |
|---|---|
Wireless headphones with Bluetooth |
2.35 |
Wired headphones |
0 |