Compare the similarity of 2 given addresses

Jonathan 380 Reputation points
2026-09-29T03:59:32.17+00:00

Hi,

The way below is good to compare the similarity of 2 given addresses? We expect to achieve that while 80% of the words (in the different position) are appearing well within both strings. Possible to achieve that? Or any other better to do that?

User's image

Developer technologies | C#
Developer technologies | C#

An object-oriented and type-safe programming language that has its roots in the C family of languages and includes support for component-oriented programming.

0 comments No comments

2 answers

Sort by: Newest
  1. Brian Pham (WICLOUD CORPORATION) 80 Reputation points Microsoft External Staff Moderator
    2026-09-29T05:38:44.76+00:00

    Hi @Jonathan ,

    The Levenshtein Distance implementation you shared is a valid approach and works well for finding small spelling differences, typos, or character-level changes between two strings.

    However, for address comparison it has some limitations. Since it compares the entire address character by character, it can produce a lower similarity score when the same address components appear in a different order

    To better align with the requirement that approximately 80% of the address words should match regardless of their position, I implemented a hybrid approach:

    public static double WordSimilarity(string first, string second)
            {
                MatchCollection firstWords = Regex.Matches(first ?? "", @"[\p{L}\p{N}]+");
                MatchCollection secondWords = Regex.Matches(second ?? "", @"[\p{L}\p{N}]+");
     
                if (firstWords.Count == 0 || secondWords.Count == 0)
                    return 0;
     
                var firstMatched = new bool[firstWords.Count];
                var secondMatched = new bool[secondWords.Count];
                int matches = 0;
     
                for (int secondIndex = 0; secondIndex < secondWords.Count; secondIndex++)
                {
                    for (int firstIndex = 0; firstIndex < firstWords.Count; firstIndex++)
                    {
                        if (!firstMatched[firstIndex] &&
                            string.Equals(firstWords[firstIndex].Value,
                                secondWords[secondIndex].Value,
                                StringComparison.OrdinalIgnoreCase))
                        {
                            firstMatched[firstIndex] = true;
                            secondMatched[secondIndex] = true;
                            matches++;
                            break;
                        }
                    }
                }
     
                for (int secondIndex = 0; secondIndex < secondWords.Count; secondIndex++)
                {
                    if (secondMatched[secondIndex])
                        continue;
     
                    string secondWord = secondWords[secondIndex].Value;
                    if (secondWord.Length < 4 || ContainsDigit(secondWord))
                        continue;
     
                    for (int firstIndex = 0; firstIndex < firstWords.Count; firstIndex++)
                    {
                        string firstWord = firstWords[firstIndex].Value;
                        if (!firstMatched[firstIndex] && firstWord.Length >= 4 &&
                            !ContainsDigit(firstWord) && WithinOneEdit(firstWord, secondWord))
                        {
                            firstMatched[firstIndex] = true;
                            matches++;
                            break;
                        }
                    }
                }
     
                return (double)matches / Math.Max(firstWords.Count, secondWords.Count);
            }
     
            public static bool IsMatch(string first, string second)
            {
                return WordSimilarity(first, second) >= 0.8;
            }
     
            private static bool ContainsDigit(string word)
            {
                foreach (char character in word)
                {
                    if (char.IsDigit(character))
                        return true;
                }
                return false;
            }
     
            private static bool WithinOneEdit(string first, string second)
            {
                if (Math.Abs(first.Length - second.Length) > 1)
                    return false;
     
                int firstIndex = 0;
                int secondIndex = 0;
                int edits = 0;
     
                while (firstIndex < first.Length && secondIndex < second.Length)
                {
                    if (char.ToUpperInvariant(first[firstIndex]) == char.ToUpperInvariant(second[secondIndex]))
                    {
                        firstIndex++;
                        secondIndex++;
                        continue;
                    }
     
                    if (++edits > 1)
                        return false;
     
                    if (first.Length >= second.Length)
                        firstIndex++;
                    if (second.Length >= first.Length)
                        secondIndex++;
                }
     
                return edits + (first.Length - firstIndex) + (second.Length - secondIndex) <= 1;
            }
        }
    

    Here's what the implementation does:

    1. Split both addresses into individual words and numbers -The regular expression extracts each address component separately instead of treating the address as one long string.
    2. Perform exact word matching first -The code compares each word from one address against the other. -Matching is case-insensitive. -The position of the words does not matter, so addresses containing the same components in different orders can still achieve a high score.
    3. Perform a second pass for unmatched words -If a word was not matched exactly, the code checks whether the words differ by only a single edit using the WithinOneEdit() method.
    4. Avoid fuzzy matching on short or numeric values

    Compared to applying Levenshtein Distance to the entire address string, I believe this is more suitable for address matching scenarios.

    If you found my response helpful or informative, I would greatly appreciate it if you could follow this guide for your confirmation.

    Thank you.

    Was this answer helpful?

    0 comments No comments

  2. Senthil kumar 2,500 Reputation points
    2026-09-29T05:20:14.99+00:00

    Hi @Jonathan

    Please try the below method will achieve the your result.

    public static double WordMatchPercentage(string a, string b)
    {
     var words1 = a.ToLower().Split(' ', ',', '.')
     .Where(x => !string.IsNullOrWhiteSpace(x))
     .ToHashSet();
     
     var words2 = b.ToLower().Split(' ', ',', '.')
     .Where(x => !string.IsNullOrWhiteSpace(x))
     .ToHashSet();
     
     int matched = words1.Intersect(words2).Count();
     
     int maxWords = Math.Max(words1.Count, words2.Count);
     
     return (double)matched / maxWords * 100;
    }
    

    Thanks.

    Was this answer helpful?


Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.