Programs without code (e.g., no executable segment) led to miss the
`env` variable, making the `CpuLoopExitPass` fail. This commit, fixes
the problem by simply populating the `cpu_loop_exit` function with a
`ret`.
In function isolation, every time we met a jump to an unexpected basic
block (i.e., a basic block that is not part of the current function), we
used to throw an exception. However this is unnecessary since oftentimes
it is sufficient to call the `function_dispatcher` or even perform a
regular function call.
The most obvious example is the case of a direct tail call. In this
situation performing a function call to the corresponding isolated
function is the most appopriate thing to do.
We used to black list memory portions target of memory read, so that
they couldn't become function entry points. However, this led to false
positives, in particular with functions that are explicitly called.
This commit ensures that we black list those addresses only if they are
not targets of function calls.
We now optimize `support.c`. We didn't do that since we were optimizing
everything after linking anyway, however, without `-O2`, functions in
`support.c` were getting an `optnone` attribute which inhibited any
optimization. In particular `newpc` and analogous functions were not
inlined. Resolving this issue led to a gigantic improvement in terms of
performance of the output program.
When optimizing a module to which function isolation has been applied the
`prune-eh` pass lead to several issues, specifically:
* `support.c` is now compiled with support for exceptions, to prevent
`raise_exception_helper` from being marked `nounwind` during the
link phase.
* Added a fake `ret` at the end of the `catchblock` to avoid promotion
of `invoke`s to regular `call`s.
* Marked the `invoke` instructions as `noinline`.
`revcc` is a simple script that forwards its arguments to a specified
compiler, except in the case in which the compiler is asked to link the
final program. In such case, the compiler is invoked as appropriate, but
the resulting binary is then translated using rev.ng and replaced by the
translated version.
We used to include `.` and `include/`, but the former is no longer
necessary and the latter has to have the highest priority to avoid that
installed headers take higher precedence, leading to a build employing
outdated headers.
This commit fixes an issue in OSRA with computations on constants. If a
constant was added to another constant, a new constant bounded value
would be created. However, such bounded value wouldn't get propagated
further. This commit fixes the problem by creating an OSRA relative to
the original constant plus an offset.
When going through basic blocks, FCI ignores basic blocks that have
already been identified as function calls. However, this lead to exclude
their fallthrough addresses from the list of fallthrough addresses.
This commit fixes this situation.
`getType` didn't have a `BlockType` to represent the entry basic block
of the `root` function. Therefore, such basic block was erroneously
identified as a translated basic block.
This commit forces the function call identification analysis during
harvesting of new basic blocks. Additionally, the appropriate jump
target reasons are associated to the involved basic blocks.
The stack analysis identifies CSV as `CPU+x` where `x` is an index that
uniquely identifies a CSV. We used to compute this index multiple times,
going through the list of global variables.
After we switched from metadata to global variables for strings
representing disassembled instructions, such process became very slow to
the point of being a bottleneck due to the large amount of global
variables.
This commit precomputes, once and for all, the unique identifier of each
CSV and saves it in a `std::set`.
`BinaryFile::readRawValue` scans the segment list to identify which
segment contains a certain address. However, it was failing if the
target address was in `.bss`, i.e., the portion of a segment after
`p_filesz` but `before `p_memsz`.
This commit lets `BinaryFile::readRawValue` return 0 in that situation.
Fix for a situation where the fallthrough basic block of an instruction
calling a helper function is not in an executable segment and therefore
not created.
This commit fixes a huge performance issue due to performing a orphan
basic block cleanup every time `JumpTargetManager::peek` was called
instead of only when actual harvesting was required.
Now using the `root` function of the cloned module as a starting point
for the isolation process.
In this way we can ignore the old module, and in particular we can
drop the `ModuleCloningVMap`, a giant map that was used to keep the
match between old and new global objects and was used in the instruction
cloning phase.
Also changed the creation of the trampoline for the isolated function,
using a new basic block and later purging all the unreachable basic
blocks in the `root` function.
Moved the definition of `OBJ` before its first use.
Moved the definition of `CSV` outside the lifting scope, to avoid
erroneous behaviors when invoking the script with the `-s` option.
This commit uses SET, information about canonical values and labels to
detect if an indirect function call is targeting an external symbol.
The strings used for the name of external symbols are uniqued global
variables. This commit also uses this approach for the disassembly of
original instructions, which used to be metadata.
So far we've been tracking only base-relative relocations in an ad-hoc
fashion. This commit introduces a data structure that can describe the
most common relocations, including those for `.got`, `.got.plt`,
base-relative and `R_*_COPY`.
A label describes a range of the binary. A label can be generated from a
symbol (basically assigning a name to range of the binary) or from a
relocation, describing the content of a certain range.
This commit generates labels from symbols and relocations, including
MIPS implicit relocations.
This commit also introduces canonical values.
A register can have a canonical value, i.e., a value that register will
assume when the analyzed module is being run. This is typically useful
for the value of the global pointer, which is different from one module
to another but, within a module, is stable.
This commit registers the canonical value of `gp` (in MIPS), if
available.
`segments_count` provides a way for the runtime to know how many
executable segments the original program had. This is used to implement
the `is_executable` function.
While the `segment_boundaries` contained only the executable segments,
`segments_count` included non-executable segments too, leading to an
out-of-bound read which sometimes led to a spurious `Unknown PC` error.
While translating the code, it might happen that a basic block needs to
be splitted and the second part to be revisited. In such cases, the
second part is register for being "purged" at the next iteration.
Translation failed if we met such a basic block before purging. This
commit correctly handles such situations.
In certain cases we find more than one instruction storing the return
address to a register. In particular, this happens with a `bltzal`
instruction in MIPS, where the return address is stored both in `ra` and
`btarget`.
For now, do not consider these as actual function calls.
We used to check if the address associated to the `PT_DYNAMIC` program
header matched the one of the `.dynamic` section. However, we were not
recording it, which is required in case sections headers are
missing/corrupt.