Memory issue due to Hash set

Jonathan 420 Reputation points
2026-10-07T04:31:40.68+00:00

Hi,

Further to handle with Hash Set, I got memory issue below

User's image

It must be due to the big set used. Here are the codes.

User's image

User's image

Developer technologies | C#
Developer technologies | C#

An object-oriented and type-safe programming language that has its roots in the C family of languages and includes support for component-oriented programming.

0 comments No comments

2 answers

Sort by: Most helpful
  1. Senthil kumar 2,585 Reputation points
    2026-10-07T05:04:50.38+00:00

    Hi @Jonathan

    addrSet = new HashSet<String>(File.Exists(@"c:/CUCCLinker/Cusfulllist.LST") ?
    File.ReadLines(@"c:/CUCCLinker/Cusfulllist.LST").Where(l => l.length>1).Select(l =>l) : Enumerable.Empty<String>(),
    StringComparer.OrdinalIgnoreCase);
    

    please check the file size may be huge. use below command check your file size.

    FileInfo fi = new FileInfo(@"c:\CUCCLinker\Custfullist.LST");
    Console.WriteLine(fi.Length / 1024 / 1024 + " MB");
    

    Thanks.

    Was this answer helpful?

    1 person found this answer helpful.

  2. Brian Pham (WICLOUD CORPORATION) 250 Reputation points Microsoft External Staff Moderator
    2026-10-07T07:31:29.4533333+00:00

    Hi @Jonathan ,

    Thank you for sharing the relevant code. I understand that the application encounters an OutOfMemoryException while processing a large collection and performing similarity comparisons.

    Based on the code, the issue is likely caused by a combination of memory usage patterns rather than the HashSet alone:

    • The HashSet<string> retains every unique line from the file in memory.
    • WordSimilarity is called once for every entry in the set. Each call creates regular-expression match collections, match objects, Boolean arrays, and temporary strings.
    • If this process runs for multiple input values, those repeated allocations can create significant memory and garbage-collection pressure.
    • The calling loop also contains a separate logic issue. Without braces or an early continue, entries containing 11 or fewer characters may be compared using the candidate from the previous iteration.

    Please update the loop as follows:

    
    foreach (string addr in addrSet)
    {
        if (addr.Length <= 11)
            continue;
    
        string candidate = addr.Substring(11);
        double score = WordSimilarity(value, candidate);
    
        if (score >= threshold && score > bestScore)
        {
            bestScore = score;
            bestMatch = addr;
            s_code2 = addr.Substring(0, 10);
            s_addr2 = addr;
        }
    }
    
    

    This loop change addresses a separate logic issue. In the original code, only the Substring assignment is controlled by the if statement. WordSimilarity is still called for entries containing 11 or fewer characters, potentially using the candidate from the previous iteration.

    Other than that, I saw your logic is:  

    var addrSet = new HashSet<string>(
        File.Exists(path)
            ? File.ReadLines(path)
                .Where(l => l.Length >= 1)
                .Select(l => l)
            : Enumerable.Empty<string>(),
        StringComparer.OrdinalIgnoreCase
    );
    
    

    But after that on the loop you only take the addr > 11 , so I think you better change to

    var addrSet = new HashSet<string>(
        File.Exists(path)
            ? File.ReadLines(path)
                .Where(l => l.Length >= 11)
                .Select(l => l)
            : Enumerable.Empty<string>(),
        StringComparer.OrdinalIgnoreCase
    );
    

    This correction may eliminate invalid or unnecessary comparisons, but it may not by itself resolve the OutOfMemoryException.

    To help determine whether the exception is related to the process architecture, the amount of data retained in memory, or repeated similarity comparisons, could you please provide the following information?

    • Project platform target: x86, x64, or Any CPU
    • Whether “Prefer 32-bit” is enabled
    • Size and approximate line count of Custfulllist.LST
    • Approximately how many input values are processed
    • The call stack from the OutOfMemoryException
    • The final value of addrSet.Count

    You can also add the following diagnostic output:

    Console.WriteLine($"64-bit process: {Environment.Is64BitProcess}");

    Console.WriteLine($"Address count: {addrSet.Count:N0}");

    Console.WriteLine($"After HashSet: {GC.GetTotalMemory(false) / 1024 / 1024:N0} MB");

     For example:

    Console.WriteLine($"Before HashSet: {GC.GetTotalMemory(false) / 1024 / 1024:N0} MB");
     
    var addrSet = new HashSet<string>(
        File.Exists(path)
            ? File.ReadLines(path)
                .Where(l => l.Length > 11)
            : Enumerable.Empty<string>(),
        StringComparer.OrdinalIgnoreCase);
     
    Console.WriteLine($"After HashSet: {GC.GetTotalMemory(false) / 1024 / 1024:N0} MB");
    Console.WriteLine($"Working set: {Environment.WorkingSet / 1024 / 1024:N0} MB");
    Console.WriteLine($"64-bit process: {Environment.Is64BitProcess}");
    Console.WriteLine($"Address count: {addrSet.Count:N0}");
    

    If the issue is reproducible, please capture a Visual Studio memory snapshot after addrSet is populated and another after processing a representative number of input values, or as close as possible to the failure.

    Comparing these snapshots can help determine whether most of the memory is retained by the hash set or whether memory usage increases significantly during the similarity comparisons.

    These details will help me identify the main source of memory consumption before recommending further code changes.

    Thanks. As the original text file is over 9GB. Any other way to do the search if it is having memory issue by Hash set?

    For this question, since the source file is over 9 GB, keeping the entire file in a HashSet<string> is likely to require a substantial amount of memory. If the current approach continues to cause OutOfMemoryException, there are several alternatives depending on whether duplicate removal is required and how frequently the file is searched.:

    1. Stream the file instead of loading it into a HashSet

    If duplicate removal is not essential, the simplest approach is to process the file line by line:

    foreach (var addr in File.ReadLines(path))
    {
        if (addr.Length <= 11)
            continue;
        string candidate = addr.Substring(11);
        double score = WordSimilarity(value, candidate);
        // ...
    }
    

    This keeps memory usage low because the entire 9+ GB file does not need to be retained in memory. The trade-off is that duplicate entries will also be processed.

    2. Use an external/disk-based data store

    If duplicate removal is important, but the dataset is too large to keep in memory, consider preprocessing the file and storing the unique records in a disk-based store such as SQLite or another database. The search can then use an appropriate index rather than keeping millions of strings in a HashSet in memory.

    3. Process the file in chunks

    Another option is to process the file in manageable batches. For example, read a portion of the data, build a temporary HashSet for that portion, perform the required processing, release the memory, and then continue with the next portion.

    This approach can preserve some of the benefits of deduplication without requiring the entire 9+ GB dataset to be resident in memory at once.

    Of these options, I would first consider streaming if duplicate removal is not a strict requirement. If deduplication is important and this file is searched repeatedly, a disk-based indexed solution would likely be more appropriate than loading the entire 9+ GB file into a HashSet.

    If you found my response helpful or informative, I would greatly appreciate it if you could follow this guide for your confirmation. Thank you.

    Was this answer helpful?


Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.