Hi @Jonathan ,
Thank you for sharing the relevant code. I understand that the application encounters an OutOfMemoryException while processing a large collection and performing similarity comparisons.
Based on the code, the issue is likely caused by a combination of memory usage patterns rather than the HashSet alone:
- The HashSet<string> retains every unique line from the file in memory.
- WordSimilarity is called once for every entry in the set. Each call creates regular-expression match collections, match objects, Boolean arrays, and temporary strings.
- If this process runs for multiple input values, those repeated allocations can create significant memory and garbage-collection pressure.
- The calling loop also contains a separate logic issue. Without braces or an early continue, entries containing 11 or fewer characters may be compared using the candidate from the previous iteration.
Please update the loop as follows:
foreach (string addr in addrSet)
{
if (addr.Length <= 11)
continue;
string candidate = addr.Substring(11);
double score = WordSimilarity(value, candidate);
if (score >= threshold && score > bestScore)
{
bestScore = score;
bestMatch = addr;
s_code2 = addr.Substring(0, 10);
s_addr2 = addr;
}
}
This loop change addresses a separate logic issue. In the original code, only the Substring assignment is controlled by the if statement. WordSimilarity is still called for entries containing 11 or fewer characters, potentially using the candidate from the previous iteration.
Other than that, I saw your logic is:
var addrSet = new HashSet<string>(
File.Exists(path)
? File.ReadLines(path)
.Where(l => l.Length >= 1)
.Select(l => l)
: Enumerable.Empty<string>(),
StringComparer.OrdinalIgnoreCase
);
But after that on the loop you only take the addr > 11 , so I think you better change to
var addrSet = new HashSet<string>(
File.Exists(path)
? File.ReadLines(path)
.Where(l => l.Length >= 11)
.Select(l => l)
: Enumerable.Empty<string>(),
StringComparer.OrdinalIgnoreCase
);
This correction may eliminate invalid or unnecessary comparisons, but it may not by itself resolve the OutOfMemoryException.
To help determine whether the exception is related to the process architecture, the amount of data retained in memory, or repeated similarity comparisons, could you please provide the following information?
- Project platform target: x86, x64, or Any CPU
- Whether “Prefer 32-bit” is enabled
- Size and approximate line count of Custfulllist.LST
- Approximately how many input values are processed
- The call stack from the OutOfMemoryException
- The final value of addrSet.Count
You can also add the following diagnostic output:
Console.WriteLine($"64-bit process: {Environment.Is64BitProcess}");
Console.WriteLine($"Address count: {addrSet.Count:N0}");
Console.WriteLine($"After HashSet: {GC.GetTotalMemory(false) / 1024 / 1024:N0} MB");
For example:
Console.WriteLine($"Before HashSet: {GC.GetTotalMemory(false) / 1024 / 1024:N0} MB");
var addrSet = new HashSet<string>(
File.Exists(path)
? File.ReadLines(path)
.Where(l => l.Length > 11)
: Enumerable.Empty<string>(),
StringComparer.OrdinalIgnoreCase);
Console.WriteLine($"After HashSet: {GC.GetTotalMemory(false) / 1024 / 1024:N0} MB");
Console.WriteLine($"Working set: {Environment.WorkingSet / 1024 / 1024:N0} MB");
Console.WriteLine($"64-bit process: {Environment.Is64BitProcess}");
Console.WriteLine($"Address count: {addrSet.Count:N0}");
If the issue is reproducible, please capture a Visual Studio memory snapshot after addrSet is populated and another after processing a representative number of input values, or as close as possible to the failure.
Comparing these snapshots can help determine whether most of the memory is retained by the hash set or whether memory usage increases significantly during the similarity comparisons.
These details will help me identify the main source of memory consumption before recommending further code changes.
Thanks. As the original text file is over 9GB. Any other way to do the search if it is having memory issue by Hash set?
For this question, since the source file is over 9 GB, keeping the entire file in a HashSet<string> is likely to require a substantial amount of memory. If the current approach continues to cause OutOfMemoryException, there are several alternatives depending on whether duplicate removal is required and how frequently the file is searched.:
1. Stream the file instead of loading it into a HashSet
If duplicate removal is not essential, the simplest approach is to process the file line by line:
foreach (var addr in File.ReadLines(path))
{
if (addr.Length <= 11)
continue;
string candidate = addr.Substring(11);
double score = WordSimilarity(value, candidate);
// ...
}
This keeps memory usage low because the entire 9+ GB file does not need to be retained in memory. The trade-off is that duplicate entries will also be processed.
2. Use an external/disk-based data store
If duplicate removal is important, but the dataset is too large to keep in memory, consider preprocessing the file and storing the unique records in a disk-based store such as SQLite or another database. The search can then use an appropriate index rather than keeping millions of strings in a HashSet in memory.
3. Process the file in chunks
Another option is to process the file in manageable batches. For example, read a portion of the data, build a temporary HashSet for that portion, perform the required processing, release the memory, and then continue with the next portion.
This approach can preserve some of the benefits of deduplication without requiring the entire 9+ GB dataset to be resident in memory at once.
Of these options, I would first consider streaming if duplicate removal is not a strict requirement. If deduplication is important and this file is searched repeatedly, a disk-based indexed solution would likely be more appropriate than loading the entire 9+ GB file into a HashSet.
If you found my response helpful or informative, I would greatly appreciate it if you could follow this guide for your confirmation. Thank you.