The symbol handling has been extended to register whether a symbol
represents a function or not. This information is then used to register,
during the global data harvesting phase, all the function symbols and
explicitly mark them through the "FunctionSymbol" `JTReason`.
We use this information during the CFEP harvesting phase to integrate
the information produced by the function boundaries detection with
potential unidentified CFEPs.
This option can be enabled with the `--use-debug-symbols`, which
supersedes `--use-sections`.
This commit introduces support for dynamic programs. The current
implementation translate the main binary and uses native libraries. This
works only if the target architecture is the same as the source
one. Currently we only handle x86-64.
* The `ExternalJumpsHandler` class has been introduced. It basically
takes care of extending the dispatcher handling the case in which the
program counter is an address outside the range of executable
addresses of the input program. In this case, a `setjmp` is perfomed,
the CPU state is serialized to physical registers and jump to the
value of the program counter is performed.
Once the target code will try to return to the translated program, a
segmentation fault will be triggered, a `longjmp` is performed and the
CPU state is deserialized so that the execution can resume (from the
dispatcher).
* `early-linked.c` has been introduced. Its purposes is to provide
declarations of variables and functions defined in `support.c`. In the
past, we had to manually create these definitions, a cumbersome and
error prone we now avoid by letting `clang` compile `early-linked.c`
and then linking it in.
* The old `support.h` is now known as `commonconstants.h`. `support.h`
now contains declarations that have to be consumed by
`early-linked.c`.
* Each architecture now provides additional information:
1. Which registers are part of the ABI and have to be preserved. If
necessary the QEMU name can be provided. For each register it's
also possible to provide their position within the `mcontext_t`
structure, provided by the signal handler.
2. Three assembly snippets, one to write a register, one to read it
and one perform an indirect jump.
Some of this information is also exposed in the output module as
metadata.
* `support.c` now installs a SIGSEGV signal handler. Since pages that
were originally executable are no longer executable, jumping there
(typically, from a library) will trigger a SIGSEGV that we will
handle. This allows us to properly deserialize the CPU state and
resume execution of the translate code.
* Now also a dynamic version of each test program is translated and
tested.
* The `merge-dynamic.py` script has been introduced: it takes case of
rewriting the translated binary so to tell the linker to performe both
the relocations of the translate program and the relocations of the
original program. It does so by rewriting a large portion of the
sections employed by the dynamic linker such as `.dynamic`, `.dynsym`
and so on.
* The `compile-time-constants.py` script has been introduced: it a
user-specified compiler on a source file producing an object
file. This object file is inspected and the value of global read-only
variables is produced in a CSV.
This simple commit should improve performance of the generated program
sensibly. Basically all the global variables will have internal linkage
from now on (unless the `--external` parameter is specified on the
command line). This way, the compiler will be able to avoid load/store
instructions when leaving code in the current translation unit.
`support.c` used to be compiled using the system compiler and then
linked to the module generated by `revamb` as a separate translation
unit. This commit introduces a change that lets `clang` compile
`support.c`. This will allow us to make the CSV static, which should
enable more aggressive optimizations.
* Change the signature of the `root` function so that it accepts an
argument: the initial value of the stack pointer, which the main is
supposed to set up. QEMU now provides us with the offset of the stack
pointer.
* Let the build system compile `support.c` for each supported
architecture, both in normal and "tracing" mode.
* Remove the `--tracing` option, this is now handled by `support.c`, in
particular depending on which version of `support.c` you link, you can
have tracing enabled or not.
* In `support.c` drop global variables representing the stack pointer,
we no longer need them.
* In `support.c` fix some warnings while handling the stack on 32-bit
architectures.
* Extende the `translate` script to handle the new way we link the final
binary and the tracing mechanism.
Introduce an option to prevent `revamb` from linking in all the QEMU
helpers. This is useful if the output doesn't need to be compiled, but
just analyzed.
This commit removes all the ELF-specific code from the `CodeGenerator`
class by creating a new class, `BinaryFile` which contains all the
information about the program that might be needed in an image format
independent way. However, `BinaryFile` has some fields which are
specific to ELF, we might want to address this when additional file
formats are supported.
A key benefit of isolating this code is that we can anticipate the
parsing of the input file, so that we have its architecture available
earlier than when `CodeGenerator` is instantiated, therefore we can drop
the `--architecture` parameter.
This commit introduces the usage of symbols, if they are available. We
employ them to produce meaningful names for basic block names.
* Collect the symbols from `.symtab`/`.dynsym`
* Box the `Segments` into a new data structure (`BinaryInfo`) which also
handles symbols.
* `JumpTargetManager::nameForAddress`: produce a meaningful name using
symbols, if possible.
* Spread some `const`-ness
* Import OSRA
* Improve the SET (aka `JumpTargetFromConstants`) by introducing the
`OperationsStack` class.
* Review `harvest` logic
* Allow to disable OSRA (along with the sumjump heuristic)
* Take the core of `getNextPC` out of it and move it to `getPC`, a
function returning both the current and the next PC. Also, fix a bug
when reaching the beginning of a basic block.
* Detect "reliable" jump targets: a "reliable" jump target is a jump
target obtained from a store to a PC but it's not a fallthrough jump.
Implement producing a CSV file containing information about the which
PCs have been translated. For each PC it is specified whether its a jump
target or not.
Instead of taking note of the executable ranges exclusively, keep track
of all the segments in `CodeGenerator`. `JumpTargetManager` instead will
keep track of executable areas only.
* Introduce the `SegmentInfo` struct, which simply holds essential
information about the segment such as start and end address,
permissions and a reference to the global variable holding its content.
* Update `CodeGenerator` to keep a vector of `SegmentInfo`.
* `JumpTargetManager`: polish the constructor and make it take the vector
of `SegmentInfo`, from which the executable ranges are then extracted.
* s/`importGlobalData`/`parseELF`/
* Save the entry point specified in the ELF header, which will be used
if the user doesn't provide an address.
* Let parse `parseELF` take care of informing libtinycode about what
has to be mmap'd and where.
* Remove some support scripts used during testing, now no longer
necessary.
* Various cleanups
Now, in `JumpTargetManager::getBlockAt`, before registering a new PC for
translation we check that the corresponding address was actually
contained in a segment marked as executable in the original binary. This
prevents translation of data, which is a problem in particular when we
will start to harvest possible code pointers from global data or
constants found in the code
* Register in `CodeGenerator::ExecutableRanges` address ranges which
contained executable code in the input ELF.
* In `JumpTargetManager::getBlockAt` check if the given PC was actually
in an executable memory area, and assert or return `nullptr` depending
on the `Try` parameter.
* Use `llvm::object` framework to obtain useful information from the ELF
binary such as pointer size and endianess.
* Introduce `CodeGenerator::importGlobalData`: import global (read-only
and writeable data) from the input binary directly into the generated
module.
* Introduce the `--linking-info` parameter: path to a CSV file where
sections containing global data extracted from the input binary are
listed with their name, start and end address.
* Expand the `Architecture` class with constructors and support accessor
methods.
* Move initialization and management of the structure describing the CPU
state (CPUStateType) into variablemanager.cpp.
* Support parts of CPU state outside "env" (e.g. the MIPSCPU
structure). Now "env" has an offset into the possibly larger CPU state
which we have to take into account where appropriate (see
VariableManager::envOffset).
* Link the helpers module into the generated module, including only what
is needed.
* Create some "no-op" or "abort" function corresponding to QEMU functions
not included in the helper module (e.g. logging and abort functions).
* Implement the CorrectCPUStateUsagePass pass, which starts from the
"env" global variable and looks for all its usages recursively, keeping
track of where pointers are pointing into the CPU state data structure,
and replaces all the load/stores with the global variable corresponding
to that specific field of the CPU state.
* After the linking phase, run SROA, the pass to adjust the CPU usage and
DCE.
* Let global variables have common linkage.