By default PerfView causes the Runtime to log an event at the beginning and end of each .NET Garbage Collection as well as every time 100K of objects are allocated. As with all events the precise time is logged, so the amount of time spent in the GC can be known. Most applications spend less than 10% of the total CPU time in the GC itself. If your application is over this %, it usually means that your allocation pattern is such that you are causing many expensive Gen 2 GCs to occur.
If the GC heap is a large percentage of the total memory used then GC heap optimization then use the Memory->Take Heap Snapshot feature to drill into GC heap usage. See Memory Usage Auditing For .NET Applications for more on memory optimization. .
During SOME GCs the application itself has to be suspended so the GC can update object references. Thus the application will pause when a GC happens. If these pause times are larger than 100ms or so, it can impact the user experience. The GC statistics will track the maximum and average pause time to allow you to confirm that this bad GC behavior is not happening.
By enabling the Finalizers option, PerfView causes the Runtime to log an event every time a managed object is finalized, meaning that its finalizer (denoted in C# with the ~ syntax) is executed. For a detailed look at the costs involved in finalization, see this blog post.
PerfView is unable to determine the stack that allocated a finalizable object, but it is able to accurately report each type of object that had a finalizer executed along with the number of instances of that type that were finalized. This data is shown on the GC Stats report in the Finalized Object Counts table.
In an ideal application implementation, all finalizable objects would be cleaned up deterministically via an object's IDisposable.Dispose implementation, which should suppress the finalization of the object. Not deterministically disposing of finalizable objects can lead to degredation of both the reliability and performance of the app. You can examine the counts reported in the Finalized Object Counts table to determine whether any types have significant numbers of instances being left for finalization. Based on that, you can examine code that creates instances of these types and determine why those instances are being left for non-deterministic cleanup rather than being cleaned up deterministically with Dispose.
PerfView tracks detailed information of what methods were Just In Time compiled. This data is mostly useful for optimizing startup (because that is when most methods get JIT compiled). If large numbers of methods are being compiled it can noticeably affect startup time. This report tells you exactly how much time is being spent (in fact exactly which methods and exactly when they were compiled). If JIT time is high, the NGen Tool can be used to precompile the code, and thus eliminate most of this overhead at startup.
As an additional option enabled with the JITInlining feature, PerfView can track all of the decisions made by the JIT about whether to inline or not at every call site. For very hot paths, the overhead of invoking methods has the potential to add measurable cost, and it's typical for developers to attempt to streamline their code as much as possible, with the intent that small methods and properties will be inlined in order to avoid these overheads. The information provided by PerfView and the JIT can be valuable in understanding when and where such attempts fail, with the JIT providing the reason it chose not to inline a particular call site, e.g. the callee had exception handling that prevented inlining, the callee was too big, the callee was explicitly annotated to prevent inlining, etc. Such information can then be used by the developer to tweak their code in pursuit of a faster outcome.
Background JIT compilation is a feature that was introduced in Version 4.5 of the .NET runtime. The basic idea is to take advantage of multiple processors available on most machines to speed up startup time by doing Just in Time (JIT) compilation on a background thread. Note that the .NET runtime's preferred solution to the cost o JIT compilation is to precompile the code with NGEN. This reduces the cost of JIT compilation to 0, where background JIT compilation cannot do nearly as well (it tends to reduce it by half), so using NGEN as part of application deployment should be considered first. However if using the NGen Tool is impossible (XCOPY deployment, non-admin deployment, silverlight, IL code generated at runtime), background JIT is the next best option.
There is a fundamental problem trying to push JIT compilation onto background threads, namely that the set of methods you will want to JIT compile depends on program execution and is not known until just before the method is used. Thus to take advantage of multiple processors you need 'oracle' that will tell you the methods you need to compile well before you will actually need to execute them.
The solution that the runtime uses is to rely on PREVIOUS runs of the same program to act as this oracle that will predict the methods that need to be compiled. For this to work the runtime will need to store on disk information about what methods were JIT compiled on the last run. Moreover, you really don't want the COMPLETE list of methods compiled because that list will include methods use well after startup and cause you to JIT compile things that are not that important. Thus to make background JIT compilation work well, we need help from the application. This is exposed in two new methods of the System.Runtime.ProfilerOptimization class introduced in .NET Version 4.5
When your code encounters a 'StartProfile' operations the runtime will do the following:
Thus by placing two simple calls in your program (typically at the beginning of Main()), you can opt into background JIT.
Background JIT has the following characteristics
It is important to realize that background JIT compilation does NOT reduce JIT time. If anything it INCREASES it, because it will JIT methods HOPING that they will be used shortly by the application. If they are not used, then that time is 'wasted'. However, the time background JIT uses is on a parallel thread (and only is attempted if there are 2 ore more processors), and thus JIT time on the background thread is effectively 'free'. Thus the important metric is how much JIT time was REMOVED from the foreground threads. As mentioned, you typically get about half, but the exact number is applications specific, and depends on how well the previous trace predicts the methods that need to be JIT compiled on this run.
If you have activated background JIT by placing the SetProfileRoot and StartProfile calls into your program you can view its effectiveness by turning on special background JIT compilation events. You do this by checking the 'Background JIT' checkbox on the advanced options of the 'Collection' dialog box. When you do this, the JITStats report is enhanced in several ways for processes that have called SetProfileRoot and StartProfile.
What can go wrong with background JIT compilation.
It your program has used SetProfileRoot and StartProfile, but JIT compilation (as shown in the JITStats view) does not show any or very little background JIT compilation, there are several issues that may be responsible. Fundamentally a important design goal was to insure that background JIT compilation did not change the behavior of the program under any circumstances. Unfortunately, it means that the algorithm tends to bail out quickly. In particular