This commit is contained in:
Elliot Killick
2025-02-25 02:24:48 -05:00
parent 4fece471ac
commit c319346761
17 changed files with 219 additions and 18 deletions
+9 -13
View File
@@ -401,7 +401,7 @@ Module constructors and destuctors are the operating system and language agnosti
In addition to module load and unload, the Windows loader may call each module's `DllMain` at `DLL_THREAD_ATTACH` and `DLL_THREAD_DETACH` times. The Windows loader only calls these routines at thread start and exit. Windows doesn't run the `DLL_THREAD_ATTACH` of a DLL following `DLL_PROCESS_ATTACH`. Additionally, a [DLL loaded after thread start won't preempt that thread to run its `DllMain` with `DLL_THREAD_ATTACH`](https://learn.microsoft.com/en-us/windows/win32/dlls/dllmain#parameters). These calls can be disabled per-library as a performance optimization by calling `DisableThreadLibraryCalls` at `DLL_PROCESS_ATTACH` time.
Compilers commonly provide access to module initialization/deinitialization functions through compiler-specific syntax. In GCC or Clang, a programmer can create module constructors/destructors using the `__attribute__((constructor))` and `__attribute__((destructor))` functions or the [`_init` and `_fini` functions, historically](https://man7.org/linux/man-pages/man3/dlopen.3.html#NOTES). Modern GCC or Clang module constructors and destructors support specifying a priority like `__attribute__((constructor(101)))` or `__attribute__((destructor(101)))` in case a particular execution order is desired.
Compilers commonly provide access to module initialization/deinitialization functions through compiler-specific syntax. In GCC or Clang, a programmer can create module constructors/destructors using the `__attribute__((constructor))` and `__attribute__((destructor))` functions or the [`_init` and `_fini` functions, historically](https://man7.org/linux/man-pages/man3/dlopen.3.html#NOTES). Modern GCC or Clang module constructors and destructors support specifying a priority like `__attribute__((constructor(101)))` or `__attribute__((destructor(101)))` (priorities of 100 and below are reserved for use by the operating system) in case a particular execution order is desired.
In C++, the constructor of an object is invoked whenever an instance of a class is created. Creating an instance of a class returns an object pointing to that instance. If an object is created in the global scope (C++ terminology) or the module scope (OS terminology) then its constructor is called during program or library initialization ([code example](code/windows/dll-init-order-test/dll-test.cpp)). If an object is created in a local scope like in a function, its constructor is called when program execution creates that object in the function. A constructor or class itself is neither inherently global nor local, it entirely depends on what context the object is created in.
@@ -411,9 +411,9 @@ Due to the useful position of constructors and destructors when run in the globa
Constructors and destructors originate from object-oriented programming (OOP), a programming paradigm first introduced by the [Simula 67](https://en.wikipedia.org/wiki/Simula) language in 1962. C++, a modern object-oriented language, was originally designed in the early 1980s as an extension of C and received initial standardization in 1998. Constructors and destructors do not exist in the C standard. On Unix systems, the concept of code that runs when a module loads and unloads goes back to the 1990 [System V Application Binary Interface Version 4](https://www.bitsavers.org/pdf/att/unix/System_V_Release_4/0-13-933706-7_Unix_System_V_Rel4_Programmers_Guide_ANSI_C_and_Programming_Support_Tools_1990.pdf) (`DT_INIT` and `DT_FINI` section types, as well as `.init` and `.fini` special section names).
In the ELF executable format, module constructors and destructors are standardized by the System V ABI to be in the `.init` and `.fini` sections. Modern systems use the non-standard but common and generally agreed-upon [`.init_array`/`.fini_array` sections](https://maskray.me/blog/2021-11-07-init-ctors-init-array), or before that the deprecated `.ctors`/`.dtors` sections. Modern GCC built binaries only include `.init_array`/`.fini_array` and `.init`/`.fini` sections, they don't include the `.ctors`/`.dtors` sections (verified with `objdump -h` and `readelf --sections`). Individually exposing each initialization/finalization routine in an array within the ELF file grants more control to the loader over calling an opaque function for handling all initialization/finalization. [A Unix-like loader loops through these routines contained in the ELF file.](https://elixir.bootlin.com/glibc/glibc-2.38/source/elf/dl-init.c#L58-L71) **Todo:** Create a code test to check whether this feature makes module initializers safely interruptible between these routines in the case of circular dependencies and make the equivalent test for Windows.
In the ELF executable format, module constructors and destructors are standardized by the System V ABI to be in the `.init` and `.fini` sections. Modern systems use the non-standard but common and generally agreed-upon [`.init_array`/`.fini_array` sections](https://maskray.me/blog/2021-11-07-init-ctors-init-array), or before that the deprecated `.ctors`/`.dtors` sections. Modern GCC built binaries only include `.init_array`/`.fini_array` and `.init`/`.fini` sections, they don't include the `.ctors`/`.dtors` sections (verified with `objdump -h` and `readelf --sections`). Individually exposing each routine in an array within the ELF file grants more control over initialization and finalization routine execution to a Unix-like loader over calling an opaque function for handling all initialization/finalization. This control and transparency lends itself to a pluggable interace that is useful in concepts such as constructor and destructor priority control ([the glibc loader does not use this per-routine knowledge to compensate for circular dependencies during module initialization](code/glibc/dlopen-init-interruption/README.md)). [A Unix-like loader loops through these routines contained in the ELF file.](https://elixir.bootlin.com/glibc/glibc-2.38/source/elf/dl-init.c#L58-L71)
The PE (Windows) executable format standard [doesn't define any sections specific to module initialization](https://learn.microsoft.com/en-us/windows/win32/debug/pe-format#special-sections); instead, a `DllMain` function or any module constructors/destructors are included with the rest of the program code in the `.text` section. MSVC optionally provides the [`init_seg`](https://learn.microsoft.com/en-us/cpp/preprocessor/init-seg) pragma to specify a section name with module constructors to run frist when compiling C++ code. However, such a section is only used if this pragma is explicitly specified by the programmer (unlikely) or in the niche cases MSVC will generate one itself (as stated by the documentation). The granularity this pragma provides is low with only `compiler`, `lib`, and `user` options. In contrast, the `.init_array`/`.fini_array` sections and `__attribute__((constructor(priority)))`/`__attribute__((destructor(priority)))` on Unix-like systems serve as a modular and robust means for controlling dynamic initializiation order.
The PE (Windows) executable format standard [does not define any sections specific to module initialization](https://learn.microsoft.com/en-us/windows/win32/debug/pe-format#special-sections); instead, a `DllMain` function or any module constructors/destructors are included with the rest of the program code in the `.text` section. MSVC optionally provides the [`init_seg`](https://learn.microsoft.com/en-us/cpp/preprocessor/init-seg) pragma to specify a section name with module constructors to run frist when compiling C++ code. However, such a section is only used if this pragma is explicitly specified by the programmer (unlikely) or in the niche cases MSVC will generate one itself (as stated by the documentation). The granularity this pragma provides is low with only `compiler`, `lib`, and `user` options. In contrast, the `.init_array`/`.fini_array` sections and `__attribute__((constructor(priority)))`/`__attribute__((destructor(priority)))` on Unix-like systems serve as a modular and robust means for controlling dynamic initializiation order.
The Windows loader calls a module's `LDR_DATA_TABLE_ENTRY.EntryPoint` at module initialization or deinitialization with the respective `fdwReason` argument (`DLL_PROCESS_ATTACH` or `DLL_PROCESS_DETACH`); it has no knowledge of `DllMain` or C++ constructors/destructors in the module scope. Merging these into one callable `EntryPoint` is the job of a compiler. For instance, [MSVC compiles a stub into your DLL (`dllmain_dispatch`) that calls any module constructors followed by `DllMain` with the `DLL_PROCESS_ATTACH` argument](code/windows/dll-init-order-test/exe-test.c) (and destructors, of course, in the reverse order). Constructors other than `DllMain`, of course, initialize in the order they are laid out in code. The word `Main` in `DllMain` indicates that `DllMain` will run as the last constructor in the module similar to how the `main` function of a program runs after all constructors. Still, I find `DllMain` to generally be a bad name because it may lead people to use constructors in ways that one might use the `main` function of a program due to the similar name (like `DllMain` is just `main` but in a DLL, which is not the case). I also find Microsoft's use of the term "entry point" (e.g. in `LDR_DATA_TABLE_ENTRY.EntryPoint`) to describe calling a module's constructor and destructor routines bad because an [entry point has a specific definition that refers to the start of program execution](https://en.wikipedia.org/wiki/Entry_point). This reason for this name stems from both an EXE and its DLLs having a `LDR_DATA_TABLE_ENTRY`. Especially since the Windows loader does accurately set the EXE's `EntryPoint` set to the program's main function (then [just above `EntryPoint` is the `DllBase` member of `LDR_DATA_TABLE_ENTRY`](https://www.geoffchappell.com/studies/windows/km/ntoskrnl/inc/api/ntldr/ldr_data_table_entry.htm), which conflates EXEs with DLLs but the other way around). So, a good question is posed by asking why the `LDR_DATA_TABLE_ENTRY` structure definition should be shared between an EXE and its DLLs at all seeing as the [as the GNU loader does not conflate these concepts](#analysis-commands.md#link_map-analysis) because, besides both being some code with data that is mapped into memory, these are completely different things. Up until one point in Windows history, [the `LDR_DATA_TABLE_ENTRY` structure definition was even shared between kernel and user-mode modules until separating into the `KLDR_DATA_TABLE_ENTRY` structure](https://www.geoffchappell.com/studies/windows/km/ntoskrnl/inc/api/ntldr/ldr_data_table_entry/index.htm): "The `LDR_DATA_TABLE_ENTRY` structure is NTDLL’s record of how a DLL is loaded into a process. In early Windows versions, this structure is similarly the kernel’s record of each module that is loaded for kernel-mode execution. The different demands of kernel and user modes eventually led to the separate definition of a `KLDR_DATA_TABLE_ENTRY`." The GNU loader calls legacy [`init` before going through the `init_array` functions](https://elixir.bootlin.com/glibc/glibc-2.38/source/elf/dl-init.c#L56) (the opposite of Windows where `DllMain` comes last after all other constructors, similar to how a `main` function would). All of these facts come together to paint a picture of Windows being too kernel-centric and monolithic, not considering the unique requirements of user-mode and correctly distinguishing between execution environments.
@@ -436,7 +436,7 @@ C# also has finalizers (historically referred to as destructors in C#). The fina
C# supports the [`ModuleInitializer` attribute](https://learn.microsoft.com/en-us/dotnet/csharp/language-reference/attributes/general#moduleinitializer-attribute) for initialization code that is to run when the assembly loads even when that assembly is a library (like traditional static constructors). Presumably, C# module initializers require protection from a global CLR initialization lock. In C# 9 and .NET 5 (released together in 2020), [module initializers were added to the language and runtime out of necessity](https://learn.microsoft.com/en-us/dotnet/csharp/language-reference/proposals/csharp-9.0/module-initializers#summary).
The unexpected initialization time of C# static constructors can cause unforseen problems similar to how Windows delay loading does for operating system initializers. For instance, a static constructor ["call is made in a locked region based on the specific type of the class"](https://learn.microsoft.com/en-us/dotnet/csharp/programming-guide/classes-and-structs/static-constructors#remarks). So, if creating an instance of a class for the first time happens at an unexpected time (perhaps via proxy through another call) like when the thread is holding a lock, and there exists a static constructor (of the same class type) that acquires the same external lock in the reverse order, then lock order inversion and consequently ABBA deadlock can occur. CLR lazy loading/initialization does have a couple significant mitigating factors that make it safer than native library lazy loading, namely: Firstly, lazy initialization can only occur upon instance creation which is necessarily more expected because it's already known that typical per-instance object constructors will run at instance creation time (unlike native library lazy loading where the initialization can potentially happen on every call to a DLL import). Though, this does leave the other, less common, static constructor trigger of referencing a member in a static class somewhat up in the air as to its safety at the given time. Secondly, static constructors are split into their own routines and initialize with granular, per-instance MT-safe synchronization instead of a broadly serializing "CLR static constructor lock", thus decreasing the chance of trying to reenter initialization or deadlocks. Lazy initialization can still become problematic if your lazy initializer routine accidentally tries to lazily initialize itself again (this issue is typically an artifact of circular dependencies). In reagard to libary loading, a synchronized, lazily initializing global type (e.g. a C# static constructor) should never load or unload libraries (or higher level .NET assemblies) to ensure that the OS loader (also CLR loader for .NET assemblies) sensibly remains at the top of the lock hierarchy. This steadfast rule must be in place to maintain lock hierarchy. If some data is only accessed from a single threaded, though, then lazy initialization may not require sychronization (synchronization is mandatory for C# static constructors and is the [the default for `Lazy<T>` types](https://learn.microsoft.com/en-us/dotnet/api/system.lazy-1?view=net-9.0#thread-safety)). Note that Microsoft documentation breaks this sensible idea on lock hierarchy by [recommending programmers call `LoadLibrary` from lazy static constructors](https://learn.microsoft.com/en-us/dotnet/csharp/programming-guide/classes-and-structs/static-constructors#usage). Regardless of synchronization, modules with significant [cross-cutting concerns](https://en.wikipedia.org/wiki/Cross-cutting_concern#Examples) should [never lazily initialize](https://devblogs.microsoft.com/oldnewthing/20070815-00/?p=25573), instead initializaing at module load-time, or preferably initializing at compile-time if possible while having little to no dependencies. From purely a performance point of view, lazy initializers could introduce ["measurable overhead"](https://learn.microsoft.com/en-us/dotnet/csharp/language-reference/proposals/csharp-9.0/module-initializers#motivation) because the language or runtime must internally perform atomic or synchronized checks to decide whether or not the initializer needs to be run on every pass (this cost is at odds with the benefit of potentially never having to run the initializer if, for example, the application goes down a different code path or errors out early... I'm looking at you, Rust `lazy_static` and `once_cell`). With all these factors in mind, it can generally be safe to use a syncrhonized, lazily initializing global type as long as an application has a clear structure that ensures a lazy initializer routine will not depend on itself through some means (directly or indirectly), and that this thinking extends to subsystems that your code depends on (e.g. the OS loader).
The unexpected initialization time of C# static constructors can cause unforseen problems similar to how Windows delay loading does for operating system initializers. For instance, a static constructor ["call is made in a locked region based on the specific type of the class"](https://learn.microsoft.com/en-us/dotnet/csharp/programming-guide/classes-and-structs/static-constructors#remarks). So, if creating an instance of a class for the first time happens at an unexpected time (perhaps via proxy through another call) like when the thread is holding a lock, and there exists a static constructor (of the same class type) that acquires the same external lock in the reverse order, then lock order inversion and consequently ABBA deadlock can occur. CLR lazy loading/initialization does have a couple significant mitigating factors that make it safer than native library lazy loading, namely: Firstly, lazy initialization can only occur upon instance creation which is necessarily more expected because it's already known that typical per-instance object constructors will run at instance creation time (unlike native library lazy loading where the initialization can potentially happen on every call to a DLL import). Though, this does leave the other, less common, static constructor trigger of referencing a member in a static class somewhat up in the air as to its safety at the given time. Secondly, static constructors are split into their own routines and initialize with granular, per-instance MT-safe synchronization instead of a broadly serializing "CLR static constructor lock", thus decreasing the chance of trying to reenter initialization or deadlocks. Lazy initialization can still become problematic if your lazy initializer routine accidentally tries to lazily initialize itself again (this issue is typically an artifact of circular dependencies). In reagard to libary loading, a synchronized, lazily initializing global type (e.g. a C# static constructor) should never load or unload libraries (or higher level .NET assemblies) to ensure that the OS loader (also CLR loader for .NET assemblies) sensibly remains at the top of the lock hierarchy. This steadfast rule must be in place to maintain lock hierarchy. If some data is only accessed from a single threaded, though, then lazy initialization may not require sychronization (synchronization is mandatory for C# static constructors and is the [the default for `Lazy<T>` types](https://learn.microsoft.com/en-us/dotnet/api/system.lazy-1?view=net-9.0#thread-safety)). Note that Microsoft documentation breaks this sensible idea on lock hierarchy by [recommending programmers call `LoadLibrary` from lazy static constructors](https://learn.microsoft.com/en-us/dotnet/csharp/programming-guide/classes-and-structs/static-constructors#usage). Regardless of synchronization, modules with significant [cross-cutting concerns](https://en.wikipedia.org/wiki/Cross-cutting_concern#Examples) should [never lazily initialize](https://devblogs.microsoft.com/oldnewthing/20070815-00/?p=25573), instead initializing at module load-time, or preferably initializing at compile-time if possible while having little to no dependencies. From purely a performance point of view, lazy initializers could introduce ["measurable overhead"](https://learn.microsoft.com/en-us/dotnet/csharp/language-reference/proposals/csharp-9.0/module-initializers#motivation) because the language or runtime must internally perform atomic or synchronized checks to decide whether or not the initializer needs to be run on every pass (this cost is at odds with the benefit of potentially never having to run the initializer if, for example, the application goes down a different code path or errors out early... I'm looking at you, Rust `lazy_static` and `once_cell`). With all these factors in mind, it can generally be safe to use a syncrhonized, lazily initializing global type as long as an application has a clear structure that ensures a lazy initializer routine will not depend on itself through some means (directly or indirectly), and that this thinking extends to subsystems that your code depends on (e.g. the OS loader).
The CLR loader, particularly the fact that it intentionally runs outside the OS loader, is a hack because only one of these two components can be at the top of the lock hierarchy and since the OS loader starts first, it should take precedence. By the CLR loader placing itself higher in the lock hierarchy than the OS loader, the CLR becomes tighly coupled with the OS loader. Ideally, the CLR under C# should be able to, as a modular subsystem, safely abstract from the OS without worrying about low-level concerns within the native loader. In particular, it should ideally be possible for C# to use the same constructors and destructors as C++ because Microsoft has tighly integrated .NET into Windows thus making it possible to accidentally utilize the technology when the programmer didn't intend to, such as via [COM interop](https://en.wikipedia.org/wiki/COM_Interop) (there are likely some cases where the Windows API internally uses .NET through COM interop in an in-process server).
@@ -468,9 +468,7 @@ The Windows loader, in contrast to Unix-like loaders, is more vulnerable to corr
- Unexpected library loading
- Inherently, delay loading may unexpectedly cause library loading when a programmer didn't intend, thus leading to [an array of potential issues that could deadlock or crash a process](#library-lazy-loading-and-lazy-linking-overview)
- MacOS previously supported lazy loading until Apple removed it, likely due to scenarios where it becomes an anti-feature and any performance gains not being worth the trade-off
- The Windows API does dynamic library loading in situations that are inappropriate in the context of an operating system
- Windows institutes that [creating a process can load libraries into the existing process](#library-loading-locations-across-operating-systems)
- The modern Universal C Runtime (UCRT) in Windows [loads libraries at process exit](code/windows/dll-process-detach-test-harness/dll-process-detach-test-harness.c)
- Windows institutes that [creating a process can load libraries into the existing process](#library-loading-locations-across-operating-systems)
- Windows runtime libraries commonly implement a poor [thread-safe](https://en.wikipedia.org/wiki/Thread_safety) implementations that restrict concurrency
- Notable runtime components in Windows such as [`atexit` registration and its callbacks](code/glibc/atexit/README.md) are not designed with deadlock-free thread safety in mind (while runtime components are not directly part of the loader, their implementations may use or integrate with it)
- Historical library loader issues
@@ -703,7 +701,7 @@ An excerpt from *Windows Internals: System architecture, processes, threads, mem
> Additionally, because lookups in linked lists are algorithmically expensive (being done in linear time), the loader also maintains two red-black trees, which are efficient binary lookup trees. The first is sorted by base address, while the second is sorted by the hash of the module’s name. With these trees, the searching algorithm can run in logarithmic time, which is significantly more efficient and greatly speeds up process-creation performance in Windows 8 and later. Additionally, as a security precaution, the root of these two trees, unlike the linked lists, is not accessible in the PEB. This makes them harder to locate by shell code, which is operating in an environment where address space layout randomization (ASLR) is enabled.
While the message on performance is a true and prudent point to make, I also find that statement alone lacks relevant perspective on the fact that `ntdll!LdrpModuleBaseAddressIndex` only exists to begin with as a workaround for Microsoft's blunder with the `LoadLibrary` function API. The point regarding security is dubious because if the module linked lists are already in the PEB, and must remain there indefinitely for backward compatibility since [Microsoft chose to share one these lists in the public `winternl.h` header](https://learn.microsoft.com/en-us/windows/win32/api/winternl/ns-winternl-peb_ldr_data), then excluding the red-black trees has no effect because security comes down to the lowest common denominator. The background information on trouble that arises from module linked lists residing in the PEB is nice (of course, there are a variety of ways to find other modules in the process but those methods would be a bit "harder" and likely not universal). Again though, there is a more relevant point to make that is not addressed by the book especially since it does cover the associated Windows history in some places, just not here.
While the message on performance is a true and prudent point to make, I also find that statement alone lacks relevant perspective on the fact that `ntdll!LdrpModuleBaseAddressIndex` only exists to begin with as a workaround for Microsoft's blunder with the `LoadLibrary` function API. The point regarding security is dubious because if the module linked lists are already in the PEB, one of which must, in practice, remain there indefinitely for backward compatibility since [Microsoft chose to share one these lists in the public `winternl.h` header](https://learn.microsoft.com/en-us/windows/win32/api/winternl/ns-winternl-peb_ldr_data), then excluding the red-black trees has no effect because security comes down to the lowest common denominator. The background information on trouble that arises from module linked lists residing in the PEB is nice (of course, there are a variety of ways to find other modules in the process but those methods would be a bit "harder" and likely not universal). Again though, there is a more relevant point to make that is not addressed by the book especially since it does cover the associated Windows history in some places, just not here.
## Investigating the Idea of MT-Safe Library Initialization
@@ -917,17 +915,15 @@ In Windows, threads are a securable resource independent of the host process:
>
> *Windows Internals: System architecture, processes, threads, memory management, and more, Part 1 (7th edition)*
Where Windows uses thread impersonation to execute code as another user, the equivalent functionality on a Unix-like system can be accomplished by forking a new process (optimized through copy-on-write memory) and using `setuid` along with the `CAP_SETUID` privilege (this is the Linux privilege, it can vary on other Unix-like OSs) to permanently change the process' user ID (UID). An OpenSSH server, for example, server works in this fashion. Process creation is fast on Unix but forking an already existing and set up process is faster and more efficient.
Where Windows often uses thread impersonation to execute code as another user, the equivalent functionality on a Unix-like system can be accomplished by creating a new minimal process and using `setuid` along with the `CAP_SETUID` privilege (this is the Linux privilege, it can vary on other Unix-like OSs) to permanently change the process' user ID (UID). An [OpenSSH server](https://isopenbsdsecu.re/mitigations/fork_exec/), for instance, works in this fashion with the project also striving towards splitting components to create [increasingly minimal processes](https://undeadly.org/cgi?action=article;sid=20241014070554). Unix is an operating system that values multiprocessing over multithreading (this is the only correct operating system design because processes contain threads). As a result, process creation is fast, making it practical to isolate each identity to its own process.
The [Principle of Least Privilege](https://en.wikipedia.org/wiki/Principle_of_least_privilege) states that user or entity should only have access to the specific data, resources, and privileges necessary to complete a required task.
Securable threads violate the Principle of Least Privilege because threads with different identities have access to each other by residing in the same address space within a process. As per *Windows Internals*, this shared access includes handles to kernel objects:
The [Principle of Least Privilege](https://en.wikipedia.org/wiki/Principle_of_least_privilege) states that user or entity should only have access to the specific data, resources, and privileges necessary to complete a required task. Securable threads violate the Principle of Least Privilege because threads with different identities have access to each other by residing in the same address space within a process. As per *Windows Internals*, this shared access includes handles to kernel objects:
> It’s important to keep in mind that all the threads in a process share the same handle table, so when a thread opens an object—even if it’s impersonating—all the threads of the process have access to the object.
Essentially, impersonation is the opposite of a proactive security design. Instead of limiting attack surface, securable threads often maximize it as much as possible.
Thread impersonation is also [highly complex](https://devblogs.microsoft.com/oldnewthing/20110928-00/?p=9533) and doesn't compose. For thread impersonation to work, every layer of a subsystem within the Windows API has to specially support its usage (e.g. [COM cloaking](https://learn.microsoft.com/en-us/windows/win32/com/cloaking)). And if there's even a single occurrence of a Windows API function being called that does not support impersonation or that you forget to pass the impersonation token to, then that's a vulnerability. Failing to correctly use and control for the consequences of thread impersonation has long been a [source of security bugs in Windows](https://y3a.github.io/2023/08/24/cve-2023-35359/) (with Microsoft now implementing hacks in the kernel to workaround this delicate security model).
Thread impersonation is also [highly complex](https://devblogs.microsoft.com/oldnewthing/20110928-00/?p=9533) and does not compose. For thread impersonation to work, every layer of a subsystem within the Windows API has to specially support its usage (e.g. [COM cloaking](https://learn.microsoft.com/en-us/windows/win32/com/cloaking)). And if there is even a single occurrence of a Windows API function being called that does not support impersonation or that you forget to pass the impersonation token to, then that creates a vulnerability. Failing to correctly use and control for the consequences of thread impersonation has long been a [source of security bugs in Windows](https://y3a.github.io/2023/08/24/cve-2023-35359/) (with Microsoft now implementing hacks in the kernel to workaround this delicate security model).
For all these reasons and more, I find that the Windows securable thread model is an insecure and fragile security model. Securable processes or per-process security is inherently more robust and secure, and this is the model that Unix-like systems are built on.
@@ -0,0 +1,10 @@
build:
$(CC) -g -shared -o lib1.so lib1.c -fPIC
# Intentionally give lib2 a circular dependency because lib1 also dynamically loads lib2
$(CC) -g -shared -o lib2.so lib2.c -fPIC -L. -l1
$(CC) -g -o main main.c -fPIC -L. -l1
clean:
rm -f main lib1.so lib2.so
.PHONY: build clean
@@ -0,0 +1,23 @@
# `dlopen` Initialization Interruption Experiment
We evaluate how the loader responds to module initialization being interrupted by a constructor in that module dynamically loading a library that holds a circular dependency on the currently initializing module.
See the [Windows equivalent](/code/windows/loadlibrary-init-circular-dependency/README.md) experiment.
## Result
```
Start: Library 1, init1_1
Start: Library 2, init2_1
Library 1, func1_1
End: Library 2, init2_1
End: Library 1, init1_1
Library 1, init1_2
Library 1, init1_3
```
The glibc loader did not go back and initialize the other constructors in library 1 before running a library 2 initializer. Technically, there is no correct behavior here because the circular dependency means taking either approach in initialization order could lead to unexpected results or crash the process. However, I would consider it "safer" for the loader to first fully initialize all the other constructors in library 1 upon calling into `dlopen` instead of initializing library 2 first. I find this approach "safer" because library 1 being fully initialized other than the one constructor calling `dlopen` seems like a probabilistically less error-prone approach than starting library 2 initialization while multiple constructors in library 1 are fully uninitialized.
Note that this discussion on the per-routine safety of module initialization interruption is only relevant to begin with because the ELF executable format distinctly exposes each routine in a module using the `.init_array` section.
**Conclusion:** The GNU loader does not utilize its knowledge of individual initialization routines to help mitigate the poor outcomes of circular dependenices on module initialization. Not trying to compensate for developer stupidity (since circular dependencies are always wrong), potentially at the cost of extra checks that could impact performance for everyone else, is a reasonable choice.
@@ -0,0 +1,31 @@
#include <stdio.h>
#include <dlfcn.h>
__attribute__((constructor))
void init1_1() {
puts("Start: Library 1, init1_1");
// Dynamically load a library with a circular dependency on the currently initializing library
void* lib = dlopen("lib2.so", RTLD_NOW);
if (!lib) {
fprintf(stderr, "%s\n", dlerror());
return;
}
puts("End: Library 1, init1_1");
}
__attribute__((constructor))
void init1_2() {
puts("Library 1, init1_2");
}
__attribute__((constructor))
void init1_3() {
puts("Library 1, init1_3");
}
void func1_1() {
puts("Library 1, func1_1");
}
@@ -0,0 +1,20 @@
#include <stdio.h>
#include <dlfcn.h>
extern void func1_1();
__attribute__((constructor))
void init2_1() {
puts("Start: Library 2, init2_1");
// The other constructors in library 1 still will not initialize before library 2 initialization even if we dynamically load library 1 explicitly
//void* lib1 = dlopen("lib1.so", RTLD_NOW);
//if (!lib1) {
// fprintf(stderr, "%s\n", dlerror());
// return;
//}
func1_1();
puts("End: Library 2, init2_1");
}
@@ -0,0 +1,3 @@
int main() {
return 0;
}
@@ -9,8 +9,8 @@ void init3() {
// Library 2 is loaded in memory, but it hasn't initialized yet
// Test what happens when we try to get a handle to the uninitialized library
//
// Note that you have to specify one of RTLD_LAZY or RTLD_NOW
// As documented in the manual, RTLD_NOLOAD can be used to promote from RTLD_LAZY to RTLD_NOW
// Note that specifying one of RTLD_LAZY or RTLD_NOW is a requirement (otherwise, we get an "Invalid argument" error)
// Specify RTLD_LAZY because, as indicated by the manual, specifying RTLD_NOW can promote the library to that binding (promotion can also occur with other flags) and we only want to test RTLD_NOLOAD
void* lib = dlopen("lib2.so", RTLD_LAZY | RTLD_NOLOAD);
puts("Still inside: Libary 3 initialization");
@@ -480,6 +480,9 @@ void startHarness() {
// The UCRT (which modern Visual Studo links programs with by default) loads the kernel.appcore.dll library at CRT exit (after running atexit handlers) messing up the last DLL in the initialization order list
// Load this DLL ahead of time to work around the issue
// Loading this library also causes RPCRT4.dll and msvcrt.dll to load, thus loading two CRTs into every UCRT process...
// The library load is done to support the ucrtbase!__acrt_AppPolicyGetProcessTerminationMethodInternal function (bloatware alert)
// - The CRT uses GetProcAddress to get the AppPolicyGetProcessTerminationMethod function from kernel.appcore.dll then calls that function
// This library loads happens after the CRT runs atexit routines and does not occur under the CRT exit lock (so at least the loader is not being placed at the bottom of a lock hierarchy here)
LoadLibrary(L"kernel.appcore.dll");
// Acquire loader lock for thread-safety:
@@ -1,5 +1,20 @@
// WORK IN PROGRESS
// We use module constructors and destructors because:
// 1. We want to avoid lazy initialization
// - Rust is a system-level programming language and we do not want our system-level code to incur the overhead that comes with constantly evaluating an atomic or synchronized "is initialized" condition as is the case with a Rust lazy_static or once_cell
// - Our subsystem should be fully initialized when it is done loading
// 2. Our module is one that presents significant cross-cutting concerns
// - As a result, we need a predictable initialization time to avoid unexpected results
// 3. Timing matters
// - Some initialization tasks require I/O, which could lead to priority inversion if they run at an unexpeted time
// - Managing priority is especially important when building a real-time application
// - Timing differences due to a lazy initialization slowdown can be used in a timing attack as a side-channel for leaking information
// 4. Two-phase initialization is an anti-pattern
// - The same can reasonably extend to lazy initialization, although lazy initialization trades a partial initialization danger for a perfomance hit by constantly checking for initialization on every use
// 5. Operating systems should be sane
// - It is not like any operating system would be broken and unhinged enough to go around terminating threads, right? Right?
pub fn hello() {
println!("Hello from the Rust dynamic library!");
}
@@ -1,6 +1,6 @@
#include <windows.h>
// Dynamically link to dummy DLL
// Dynamically link to DLL
__declspec(dllimport) void DummyExport();
int main() {
@@ -1,6 +1,6 @@
#include <windows.h>
// Dynamically link to dummy DLL
// Dynamically link to DLL
__declspec(dllimport) void DummyExport();
int main() {
@@ -0,0 +1,17 @@
# `LoadLibrary` Initialization Circular Dependency Experiment
We evaluate the loader's behavior when encountering a circular dependency during module initialization.
See the [glibc equivalent](/code/glibc/dlopen-init-interruption/README.md) experiment.
## Result
```
Start: DLL 1 Init
Start: DLL 2 Init
DLL 1, func
End: DLL 2 Init
End: DLL 1 Init
```
**Conclusion:** As expected, the Windows loader has no choice but to initialize DLL 2 while DLL 1 could be partially initialized due to a circular dependency. The PE executable format exposes all initialization and finalization tasks as merged into one externally callable `EntryPoint` and arbitrary code cannot be interrupted or preempted, so the Windows loader's current behavior is the best it can be under the conditions.
@@ -0,0 +1,19 @@
@echo off
rem These options replicate Visual Studio
rem /MD: Use UCRT instead of statically linking CRT into modules
rem /INCREMENTAL:NO: Remove "ILT" from symbol names
cl lib1.c /DUNICODE /D_UNICODE /MD /LD /Zi /DEBUG /link /INCREMENTAL:NO
rem Intentionally give lib2 a circular dependency because lib1 also dynamically loads lib2
cl lib2.c lib1.lib /DUNICODE /D_UNICODE /MD /LD /Zi /DEBUG /link /INCREMENTAL:NO
cl exe-test.c lib1.lib /DUNICODE /D_UNICODE /MD /Zi /DEBUG /link /INCREMENTAL:NO
rem Helper: If compilation fails because cl command doesn't exist then re-run in the correct environment
if %ERRORLEVEL% equ 9009 (
echo Compiler not found. Opening Visual Studio developer environment...
set "VSCMD_START_DIR=%cd%"
set "VSCMD_SKIP_SENDTELEMETRY=1"
rem Find the latest x64 Developer CMD version then open it
rem ^&: Inject running our build script after Visual Studio script (bug -> feature)
for /f "delims=" %%f in ('dir /b /s /A-D /O-N "%PROGRAMDATA%\Microsoft\Windows\Start Menu\Programs\x64 Native Tools Command Prompt for VS *.lnk"') do start "" "%%f" ^& build.bat & goto :EOF
)
@@ -0,0 +1,11 @@
#include <windows.h>
// Dynamically link to test DLL
__declspec(dllimport) void DummyExport();
int main() {
// Call dummy export so our link to the DLL isn't optimized out
DummyExport();
return EXIT_SUCCESS;
}
@@ -0,0 +1,28 @@
#include <windows.h>
#include <stdio.h>
EXTERN_C __declspec(dllexport) void func() {
puts("DLL 1, func");
}
BOOL WINAPI DllMain(HINSTANCE hinstDLL, DWORD fdwReason, LPVOID lpReserved) {
switch (fdwReason) {
case DLL_PROCESS_ATTACH:
puts("Start: DLL 1 Init");
HMODULE lib2 = LoadLibrary(L"lib2.dll");
if (!lib2) {
__debugbreak();
}
puts("End: DLL 1 Init");
break;
}
return TRUE;
}
EXTERN_C __declspec(dllexport) void DummyExport() {
// Exported function that does nothing so we can dynamically link
//puts("Test export");
}
@@ -0,0 +1,25 @@
#include <windows.h>
#include <stdio.h>
// Dynamically link to func from lib1
__declspec(dllimport) void func();
BOOL WINAPI DllMain(HINSTANCE hinstDLL, DWORD fdwReason, LPVOID lpReserved) {
switch (fdwReason) {
case DLL_PROCESS_ATTACH:
puts("Start: DLL 2 Init");
// Creating the circular dependency via dynamic linking or dynamic loading (or doubling down by doing both) gives the same result, which essentially means this LoadLibrary is a no-op because library 1 is already in its initialization phase (and code cannot arbitrarily be interrupted or preempted)
//HMODULE lib1 = LoadLibrary(L"lib1.dll");
//if (!lib1) {
// __debugbreak();
//}
func();
puts("End: DLL 2 Init");
break;
}
return TRUE;
}
@@ -1,7 +1,7 @@
#include <windows.h>
#include <stdio.h>
// Dynamically link to dummy DLL
// Dynamically link to DLL
__declspec(dllimport) void DummyExport();
#define NUM_THREADS 10000