A high-level, general-purpose programming language, created as an extension of the C programming language, that has object-oriented, generic, and functional features in addition to facilities for low-level memory manipulation.
Hi @Batorek
Short version first: no C++ standard, including C++98 (ISO/IEC 14882:1998), gives you a way to pin a function into the processor cache. The CPU decides what lives in cache based on what's executed recently and how close the code is in memory. What you can do is write and build the code so the hot functions stay small and sit next to each other, which is what keeps them resident in practice.
The approach that gives the biggest effect in Visual C++:
- Keep hot functions compact: no large local arrays, few branches, no exceptions in the hot path, and loops that walk memory sequentially.
- Group them in memory with
__declspec(code_seg("hot"))on each hot function so the linker places them together in one section, and mark the small ones__forceinlineso calls disappear. - Build with /O2 /GL and link with /LTCG, then use Profile-Guided Optimization (/GENPROFILE, run a representative workload, then /USEPROFILE). PGO reorders your functions and basic blocks by actual execution frequency, which is the closest thing to "cache-aware layout" the toolchain offers.
- For the data those functions touch, keep it contiguous and aligned (
__declspec(align(64))) so each cache line carries useful bytes.
Are you targeting a specific CPU, and is your concern the instruction cache or the data your functions read? The answer changes a bit depending on which one is the bottleneck.
If this helped, please click Accept Answer so others with the same question can find it.
References: https://learn.microsoft.com/en-us/cpp/build/profile-guided-optimizations https://learn.microsoft.com/en-us/cpp/cpp/code-seg-declspec