This commit is contained in:
Elliot Killick
2025-02-20 00:33:58 -05:00
parent f4e3fd21ab
commit b347098982
+24 -17
View File
@@ -57,9 +57,10 @@ All of the information contained here covers Windows 10 22H2 and glibc 2.38 on L
- [Synchronization Requirements](#synchronization-requirements)
- [Flimsy Thread-Local Data](#flimsy-thread-local-data)
- [The PEB Problem](#the-peb-problem)
- [Enforces Dynamic Initializatio](#enforces-dynamic-initializatio)
- [Enforces Dynamic Initialization](#enforces-dynamic-initialization)
- [Adds Process Startup Overhead](#adds-process-startup-overhead)
- [Promotes Centralization](#promotes-centralization)
- [Elevates Backward Compatability Risk](#elevates-backward-compatability-risk)
- [Weakens Security](#weakens-security)
- [Summary](#summary-2)
- [Procedure/Symbol Lookup Comparison (Windows `GetProcAddress` vs POSIX `dlsym` GNU Implementation)](#proceduresymbol-lookup-comparison-windows-getprocaddress-vs-posix-dlsym-gnu-implementation)
@@ -458,11 +459,11 @@ The Windows loader, in contrast to Unix-like loaders, is more vulnerable to corr
- Windows kernel mode and user mode closely integrate (NT and NTDLL), whereas [Unix began with modularity as a core value](https://en.wikipedia.org/wiki/Unix_philosophy)
- This value carried through to the formalization of Unix in the POSIX and C standards, and the System V ABI specification
- Windows overrelies on dynamic initialization and dynamic operations in general
- It is always best practice for robustness and performance to initialize statically (i.e. at compile time) over dynamically (using module initializers and finalizers including Windows `DllMain`) if possible
- It is always best practice for robustness and performance to initialize statically (i.e. at compile time) over dynamically (using module initializers and finalizers including Windows `DllMain`) if feasible
- Windows commonly requires dynamic initialization even for core system functionality, such as [initializing a critical section](https://learn.microsoft.com/en-us/windows/win32/api/synchapi/nf-synchapi-initializecriticalsection)
- The Windows Process Environment Block (PEB), along with its common use by getter functions like `GetProcessHeap`, artificially [enforces dynamic initialization](#the-peb-problem)
- The multithreading-first design of Windows can increase contention when accessing popular shared resources such as the process heap, which may require some Windows components [dynamically create their own resources](https://web.archive.org/web/20140805104223/https://blogs.msdn.com/b/oleglv/archive/2003/10/28/56142.aspx#:~:text=static%20CRT%20allocates%20its%20heap%20in%20DllMain%20of%20the%20owning%20DLL)
- In contrast, POSIX data structures commonly provide a [static initialization option](#constructors-and-destructors-overview) and Unix prioritizes multiprocessing through pluggable processes
- In contrast, POSIX data structures commonly provide a [static initialization option](#constructors-and-destructors-overview), and Unix prioritizes multiprocessing through minimal and pluggable processes
- Unexpected library loading
- Inherently, delay loading may unexpectedly cause library loading when a programmer didn't intend, thus leading to [an array of potential issues that could deadlock or crash a process](#library-lazy-loading-and-lazy-linking-overview)
- MacOS previously supported lazy loading until Apple removed it, likely due to scenarios where it becomes an anti-feature and any performance gains not being worth the trade-off
@@ -699,7 +700,7 @@ An excerpt from *Windows Internals: System architecture, processes, threads, mem
> Additionally, because lookups in linked lists are algorithmically expensive (being done in linear time), the loader also maintains two red-black trees, which are efficient binary lookup trees. The first is sorted by base address, while the second is sorted by the hash of the module’s name. With these trees, the searching algorithm can run in logarithmic time, which is significantly more efficient and greatly speeds up process-creation performance in Windows 8 and later. Additionally, as a security precaution, the root of these two trees, unlike the linked lists, is not accessible in the PEB. This makes them harder to locate by shell code, which is operating in an environment where address space layout randomization (ASLR) is enabled.
While the message on performance is a true and prudent point to make, I also find that statement alone lacks relevant perspective on the fact that `ntdll!LdrpModuleBaseAddressIndex` only exists to begin with as a workaround for Microsoft's blunder with the `LoadLibrary` function API. The point regarding security is dubious because if the module linked lists are already in the PEB (and must remain there indefinitely for backward compatibility) then excluding the red-black trees has no effect because security comes down to the lowest common denominator. The background information on trouble that arises from module linked lists residing in the PEB is nice (of course, there are a variety of ways to find other modules in the process but those methods would be a bit "harder" and likely not universal). Again though, there is a more relevant point to make that is not addressed by the book especially since it does cover the associated Windows history in some places, just not here.
While the message on performance is a true and prudent point to make, I also find that statement alone lacks relevant perspective on the fact that `ntdll!LdrpModuleBaseAddressIndex` only exists to begin with as a workaround for Microsoft's blunder with the `LoadLibrary` function API. The point regarding security is dubious because if the module linked lists are already in the PEB, and must remain there indefinitely for backward compatibility since [Microsoft chose to share one these lists in the public `winternl.h` header](https://learn.microsoft.com/en-us/windows/win32/api/winternl/ns-winternl-peb_ldr_data), then excluding the red-black trees has no effect because security comes down to the lowest common denominator. The background information on trouble that arises from module linked lists residing in the PEB is nice (of course, there are a variety of ways to find other modules in the process but those methods would be a bit "harder" and likely not universal). Again though, there is a more relevant point to make that is not addressed by the book especially since it does cover the associated Windows history in some places, just not here.
## Investigating the Idea of MT-Safe Library Initialization
@@ -940,7 +941,7 @@ Let's get some numbers on Windows vs. Linux thread creation and join times for 1
**Benchmark systems details:** Both Xen HVMs, Intel i5 4590, 4 vCPUs each, 8 GiBs of memory each, up-to-date Windows 10 22H2 and Fedora 39 on Linux 6.1. Tests performed while the host system and other virtual machines were suspended or turned off.
Linux native thread creation and join times comes out firmly ahead, averaging speeds 3.2x faster than Windows. However, outside of some server applications that may correspond each client connection to a new thread, quickly creating 10,000 threads is not a realistic workload. Upon booting Windows, Process Explorer shows that there are about 1,000 threads between all processes on the system. So, Windows thread creation time is unlikely to become a performance bottleneck in practice especially because Windows typically keeps threads alive and waiting as worker threads for some time instead of immediately deleting them in case new work comes along. Windows threads run their `DLL_THREAD_ATTACH` and `DLL_THREAD_DETACH` routines at thread startup and exit, which requires the same synchronization as `LoadLibrary` (including the `DLL_PROCESS_ATTACH` routine) and `FreeLibrary` (including the `DLL_PROCESS_DETACH` routine) operations. Therefore, significant variance or unexpected stutters could be present in thread startup and exit times if overlapping thread creation/exit or library load/free operations occur. Interestingly, by looking at these numbers we can see that Windows thread creation overhead primarily comes from the time it takes for the NT kernel to spawn the thread itself and not from any action in user-mode (in the future, we may run tests to see how perfomance changes with two threads simultaneously creating threads since the synchronization requirement of thread loader initialization could have a greater effect then). Another minor note is that not joining threads on Linux yields around a 25% performance improvement, although this didn't seem to have a noticeable effect on Windows (joining means running the thread until its end so any perfomance impact here would be due to the scheduler).
Linux native thread creation and join times comes out firmly ahead, averaging speeds 3.2x faster than Windows. However, outside of some server applications that may correspond each client connection to a new thread, quickly creating 10,000 threads is not a realistic workload. Upon booting Windows, Process Explorer shows that there are about 1,000 threads between all processes on the system. So, Windows thread creation time is unlikely to become a performance bottleneck in practice especially because Windows typically keeps threads alive and waiting as worker threads for some time instead of immediately deleting them in case new work comes along. Windows threads run their `DLL_THREAD_ATTACH` and `DLL_THREAD_DETACH` routines at thread startup and exit, which requires the same synchronization as `LoadLibrary` (including the `DLL_PROCESS_ATTACH` routine) and `FreeLibrary` (including the `DLL_PROCESS_DETACH` routine) operations. Therefore, significant variance or unexpected stutters could be present in thread startup and exit times if overlapping thread creation/exit or library load/free operations occur. Interestingly, by looking at these numbers we can see that Windows thread creation overhead primarily comes from the time it takes for the NT kernel to spawn the thread itself and not from any action in user-mode (in the future, we may run tests to see how performance changes with two threads simultaneously creating threads since the synchronization requirement of thread loader initialization could have a greater effect then). Another minor note is that not joining threads on Linux yields around a 25% performance improvement, although this didn't seem to have a noticeable effect on Windows (joining means running the thread until its end so any performance impact here would be due to the scheduler).
Next, we will review the difference in resource consumption between Windows and Unix threads. Each thread requires its own stack memory allocation. On Windows, the default reservation size of this memory mapping is [1 MiB](https://learn.microsoft.com/en-us/windows/win32/procthread/thread-stack-size). However, only [64 KiBs](https://devblogs.microsoft.com/oldnewthing/20031008-00/?p=42223) of that reservation is consumed from phyiscal memory. On Linux, these sizes are [8 MiB](https://unix.stackexchange.com/a/473445) and the [archiecture's page size](https://stackoverflow.com/a/24819521) (typically 4 KiBs on modern x86-based and ARM systems), respectively. Since Linux and other Unix-like systems follow the system's page size (the smallest possible memory mapping size as set by the MMU) when creating memory mappings, each thread's stack memory mapping consumes significantly less memory on Unix systems than on Windows, an attribute which is certainly desirable for a general-purpose computer. Specifically, with a page size of 4 KiB means Linux threads are 16x more lightweight in memory than Windows. These facts only account for user-mode threads because kernel-mode threads do not exist in virtual memory. Kernel-mode threads are fully committed into physical memory with a fixed size stack (typically [8 KiBs on Linux](https://docs.kernel.org/next/x86/kernel-stacks.html) or [16 KiBs on Windows](https://bsodtutorials.blogspot.com/2013/11/kernel-stacks-user-stacks-dpc-stacks.html)). Generally, less memory mapping granularity also makes guard pages less effective at catching memory overrun bugs.
@@ -950,7 +951,7 @@ With more threads, in particular running threads, comes greater context switchin
Checking in Performance Monitor, a freshly booted Windows system does around 700 context switches per second while idling. In contrast, checking `vmstat` shows that an idling desktop Linux system sees around 150 context switches per second. Linux, probably due to less overall background work going on, has around 5 times less context switches while idling. These differences are significant and could lead to notable baseline performance and battery life differences (although, Windows may takes steps to reduce background work when running on a battery for devices like laptops).
A kernel's scheduler plays a large role in the perfomance of a multithreading program or threads within a system. A scheduler decides which thread the kernel should switch to when a context switch occurs. As of version 6.6 (2023), the Linux kernel uses a new scheduler called [earliest eligible virtual deadline first (EEVDF)](https://en.wikipedia.org/wiki/Earliest_eligible_virtual_deadline_first_scheduling) that employs an algorithm by the same name. This scheduler takes into account multiple paramters including virtual time, eligible time, virtual requests and virtual deadlines for determining scheduling priority. This scheduler replaces the Completely Fair Scheduler (CFS), likely because overly striving for equal run-time distribution over thread readiness factors can cause [lock convoys](https://learn.microsoft.com/en-us/archive/msdn-magazine/2008/october/concurrency-hazards-solving-problems-in-your-multithreaded-code#lock-convoys) in practice. According to *Windows Internals: System architecture, processes, threads, memory management, and more, Part 1 (7th edition)*, "Windows implements a priority-driven, preemptive scheduling system". Indeed, the Windows scheduler is dynamic, relying heavily on ["priority boosts"](https://learn.microsoft.com/en-us/windows/win32/procthread/scheduling) to optimize for the foreground window, user interface interactiveness, multimedia applications, and lock ownership for locks that fully rely on the kernel (e.g. an event object allows execution to proceed). The *Windows Internals* book talks in-depth about the scheduler in its "Thread scheduling" subchapter. The scheduler on Windows and Linux is best optimized for the workloads common to each system (similar to a heap memory allocator, there are no settings that will be best optimized for every possible workload).
A kernel's scheduler plays a large role in the performance of a multithreading program or threads within a system. A scheduler decides which thread the kernel should switch to when a context switch occurs. As of version 6.6 (2023), the Linux kernel uses a new scheduler called [earliest eligible virtual deadline first (EEVDF)](https://en.wikipedia.org/wiki/Earliest_eligible_virtual_deadline_first_scheduling) that employs an algorithm by the same name. This scheduler takes into account multiple paramters including virtual time, eligible time, virtual requests and virtual deadlines for determining scheduling priority. This scheduler replaces the Completely Fair Scheduler (CFS), likely because overly striving for equal run-time distribution over thread readiness factors can cause [lock convoys](https://learn.microsoft.com/en-us/archive/msdn-magazine/2008/october/concurrency-hazards-solving-problems-in-your-multithreaded-code#lock-convoys) in practice. According to *Windows Internals: System architecture, processes, threads, memory management, and more, Part 1 (7th edition)*, "Windows implements a priority-driven, preemptive scheduling system". Indeed, the Windows scheduler is dynamic, relying heavily on ["priority boosts"](https://learn.microsoft.com/en-us/windows/win32/procthread/scheduling) to optimize for the foreground window, user interface interactiveness, multimedia applications, and lock ownership for locks that fully rely on the kernel (e.g. an event object allows execution to proceed). The *Windows Internals* book talks in-depth about the scheduler in its "Thread scheduling" subchapter. The scheduler on Windows and Linux is best optimized for the workloads common to each system (similar to a heap memory allocator, there are no settings that will be best optimized for every possible workload).
Departing from synthetic benchmarks, [Linux is also better equipped to take advantage of modern CPUs with high core counts in real-world applications](https://www.phoronix.com/review/3990x-windows-linux/6), with increasing margins for higher numbers of cores.
@@ -1036,23 +1037,29 @@ Thread-local data is a fragile mechanism that introduces unnecessary failure poi
**!!! WORK IN PROGRESS !!!**
In Windows, the Process Environment Block (PEB) exists to store data that is global to the process. As opposed to dynamic linking with data symbols made available by libraries, the PEB is a bad model for a variety of reasons:
In Windows, the Process Environment Block (PEB) stores data that is global to the process. As opposed to linking with data symbols made available by libraries, the PEB is a bad model for several reasons:
### Enforces Dynamic Initializatio
### Enforces Dynamic Initialization
Reducing module dynamic init can reduce the harm circular dependnenices has on libraries.
Instead of resolving data references at link time, the PEB, and its various getter functions such as `GetProcessHeap`, force developers to initialize dynamically via the module constructor. Insisting on dynamic initialization when none should be necessary is suboptimal for performance. In addition, while dependency cycles should never exist as a core principle, reducing or eliminating a module's reliance on dynamic initialization can work as a mitigating factor for preventing crashes due to these loops, should they crop up.
getprocessheap getlasterror
**TODO:** Create programs testing dynamic linking of data symbols with Windows IAT vs glibc symbol table
Windows IAT also slow because it requires double dereferencing each time a data symbol is used.
### Adds Process Startup Overhead
PEB structure must dynamically init as well as forcing modules to dynamically initialize. Two-sided or double dynamic initialization.
Dynamic linking would only add overhead for the necessary data symbols. Windows IAT also slow because it requires double dereferencing each time a data symbol is used.
Each process must initialize its PEB, a large and complex data structure, on every startup. This inefficient burden contributes to the slow process creation times on Windows. Since the PEB also enforces dynamic initialization on the side of modules, the PEB essentially gives way to two-sided or double dynamic initialization.
### Promotes Centralization
Promotes monolithic design instead of dlls exposing data symbols
By its very definition, the PEB centralizes state into one process-wide block. The contents of this block are dictated only by Microsoft. Clumping together unrelated data like this makes the operating system more monolithic as opposed to an improved way of functioning whereby libraries, each of which is its own subsystem, simply choose data symbols to export.
Centralizing state is also bad because it encourages coarse-grained locking, increasing lock contention, as is true for the PEB with its single `FastPEBLock` for synchronizing access even to unrelated members in this data structure. Windows calls this critical section lock "fast" presumably because it has a spin count attached to it that optimizes for the typically short acquisition times of this lock. However, a heuristic that improves performance by busy waiting is not a better solution than implementing fine-grained locking, thus reducing waits to begin with. Common scenarios where acquiring the PEB lock is necessary includes [creating thread-local data](https://github.com/reactos/reactos/blob/54433319af31c2b49737469d36072153de375f4d/dll/win32/kernel32/client/thread.c#L1109) and accessing environment variables, whereas Unix systems typically use purpose-built locks or low-level atomic operations, even including MT-safe synchronization, in these cases.
### Elevates Backward Compatability Risk
The PEB's definition is well-known and Microsoft must stay compatible with its exact layout through Windows versions. In contrast, data symbols do not require positioning into any exact layout and can easily be versioned.
### Weakens Security
@@ -1060,7 +1067,7 @@ Through a thinly veiled segment register, the PEB exposes pointers to various va
### Summary
The PEB should not exist
The PEB should never have existed. Despite Microsoft [first adding the PEB in Windows NT 3.1](https://www.geoffchappell.com/studies/windows/km/ntoskrnl/inc/api/pebteb/peb/index.htm#Layout:~:text=3.10) (the first Windows NT release), which is also when the operating system first gained the ability for dynamic linking, how the PEB works is a callback to the times before dynamic linking, when data was centralized into one location and required manual symbol resolution to obtain. Today, the PEB still exists and Microsoft is not shy about expanding it, undermining dynamic linking with each new member it receives.
## Procedure/Symbol Lookup Comparison (Windows `GetProcAddress` vs POSIX `dlsym` GNU Implementation)
@@ -1318,7 +1325,7 @@ The modern loader explicitly checks `ntdll!LdrInitState` to optionally perform l
Be aware that for both the legacy and modern loaders, this improvement in startup performance comes with a trade-off in run-time performance. Since, after loader initialization is complete those branches on `ntdll!LdrpInLdrInit` or `ntdll!LdrInitState` become nothing but dead weight.
Beyond a slight startup perfomance improvement through reduced synchronization overhead, there is another reason Microsoft may want to disallow new (potential load owner) threads from running during process statup, particularly in regard to module initializers: Windows places loader lock (or the modern equivalents) at the bottom of any lock hierarchy that is external to the loader. Thus, restricting additional threads from running in this stage works as a quick fix to mitigate ABBA deadlock risk from module initializers at process startup. Of course, this risk reduction does not extend to libraries that are dynamically loaded at process run-time.
Beyond a slight startup performance improvement through reduced synchronization overhead, there is another reason Microsoft may want to disallow new (potential load owner) threads from running during process statup, particularly in regard to module initializers: Windows places loader lock (or the modern equivalents) at the bottom of any lock hierarchy that is external to the loader. Thus, restricting additional threads from running in this stage works as a quick fix to mitigate ABBA deadlock risk from module initializers at process startup. Of course, this risk reduction does not extend to libraries that are dynamically loaded at process run-time.
The GNU loader doesn't implement any such startup performance hack to forgo locking on process startup. The absence of any such mechanism by the GNU loader enables threads to start and exit at process startup or within a module initializer. The same is true for process exit. Therefore, the GNU loader is more flexible in this respect.
@@ -1905,7 +1912,7 @@ I've provided Microsoft lots of constructive criticism and ways to fix their sys
- Complaint #9: [PowerShell Null Comparison "Design"](https://learn.microsoft.com/en-us/powershell/utility-modules/psscriptanalyzer/rules/possibleincorrectcomparisonwithnull)
- Hey, this is how it "works by-design", the PowerShell designers really went back to the drawing board to engineer some state of the art stuff here (nothing against the PowerShell devs personally, of course, but you're making this too easy for me)
- This one always gets a laugh out of me
- To be fair, this work could be the product of so-called Microsoft double agents (or just cool guys) intentionally inserting laughably poor design into Windows, in which case, good job guys you have really out done yourselves this time
- To be fair, this masterpiece could be from the goated Microsoft employees who always come through to slip some hilariously bad design into Windows. In that case, bravo—you’ve really outdone yourselves this time!
- [Complaint #10](https://elliotonsecurity.com/perfect-dll-hijacking/shellexecute-initial-deadlock-point-stack-trace.png)
- Complaint #11: [Forced Recall](https://www.theverge.com/2024/9/2/24233992/microsoft-recall-windows-11-uninstall-feature-bug)
- Nobody wants Recall, literally noone