This commit simply reduces the length of lines in CMake by splitting the
statements over multiple lines. This is particularly useful when listing
the translation units composing a program/library. In fact, it makes
merge much easier.
`enum`s throw their entries into the parent scope. `enum class` can work
around this issue, however, `enum class` cannot have methods. Since we
often need to have `getName` and `fromName` functions for the enums, we
now create a namespace for each enum that contain a number of helper
functions.
Doxygen doc used to look for source files in the source root directory
only. We now use `git ls-files` to figure out which files need to be
part of the documentation.
The `CounterMap` class allows to increment a counter associated to an
arbitrary key. This is useful for printing out statistics, e.g., per
`Function` or `BasicBlock`.
`.clang-format` is the configuration file for the `clang-format` tool,
which can help us to enforce a consistent coding style.
`check-conventions.sh` is a a simple bash script that checks (using
mostly `git grep`) if, after running `clang-format`, there are some
undesired situations such as lines ending with "(" or "<".
This script should produce no output before a merge request is
merged. However, currently this is not the case, therefore the script
should be used mostly for new code (in particular, files) only.
A set of assertion-related functions has been introduced:
* `revng_abort(message)`: aborts, in release builds too.
* `revng_check(what, message)`: asserts `what`, in release builds
too. Also emits a `__builtin_assume`, that can lead to additional
optimizations in clang.
* `revng_unreahcable(message)`: identical to `revng_abort`, but in
release builds emits `__built_unreachable`.
* `revng_assert(what, message)`: asserts in debug builds, otherwise
emits `sizeof(what)` (to suppress unused variable warnings) and
`__builtin_assume`.
The adoption of these function has the following benefits:
* Nice stack traces.
* The developer can choose to enforce an `assert` (or an `unreachable`)
at release-time too by using `check`/`abort`.
* Most warnings about unused variables in release mode should be gone.
* When using clang, the `assert`s become `assume`s, which might enable
additional optimizations (with no run-time costs).
* The `assert(Condition && "Reason")` trick is no longer needed, we now
have a proper argument.
Equality comparisons used to be ignored more often than required due to
the fact that they give no hint on the signedness of the tracked value.
This commits removes some assertions and improves the handling of OSRs
without a known signedness. In particular, now, inequalities can be
solved even in absence of signedness information, as long as the result
that you would get with a signed OSR and an unsigned OSR matches (i.e.,
the signedness doesn't matter).
We used to have a special handling of unsigned comparisons, since we
assumed that each side of the comparison had to be greater than or equal
to zero. This commit further widens the cases in which this is
useful. Specifically, if both the comparison we're dealing with and the
greater-than-or-equal-to-zero comparison don't have an upper bound, we
flip one of the two in a way that ensure that they represent a closed
interval. Then, if the flipped comparison is the former, we reflip the
final result.
If a conditional branch propagates a constraint on a value that is
exactly the opposite with respect to its current constraint, we simply
ignore it, since simply flipping the condition (and the destination
basic blocks) would do the same.
In future, we should propagate a contradiction on the appropriate
branch.
We used to ignore the `.gnu.hash` section, however it turns out to be
fundamental in case the ELF we're working on *defines* one or more
symbols. This has not been a problem so far since we usually work with
the main executable only, which, usually, doesn't define any dynamic
symbol. However, for example, `ls` "defines" a `getoptind` symbol (more
accurately, it clones it from libc).
To handle this situation, we replace the `.gnu.hash` with a minimal
`.hash` section. Basically, the `.hash` section should contain an hash
table. However, currently, we just create an hash table with a single
entry, which immediately triggers scanning the chain of entries
colliding in the (only) entry.
The `FunctionCallIdentification` pass now has an API to get the
fallthrough basic block of a function call basic block and to check
whether a certain basic block/address is the fallthrough of a function
call.
This commit introduces a `dump` method for `OSR` and `BoundedValue` for
easier debugging within gdb. It also introduces `debug_function`, a
definition that wraps attributes to ensure the function is emitted even
if unused and emitted as a standalone function that can be called from
gdb.
`BoundsIterator` had an issue in case there was the need to iterat over
bounds reaching the upper bound of an integer: the increment would
generate an overflow that would go undetected and, therefore, lead to an
infinite loop.
This commit introduces support for dynamic objects. We do not support
translating dynamic libraries yet, therefore this commit introduces
support for PIE programs.
At the current stage, QEMU does not provide us explicit information
about an instruction using the program counter, but introduces its value
as an immediate. As a consequence, we cannot support arbitrary
relocation. For this reason, we statically relocate the program to a
fixed address (`0x50000000` by default, but it can be customized through
the `--base` argument). Therefore, all the addresses read from ELF data
structure need to be relocated.
Code compiled with `-fPIC` cannot store in global data the address of a
function, since it will be relocated at run-time. This means that the
global data harvesting won't bring any benefit. On the other hand, going
through dynamic symbols can be hugely beneficial. Same argument for
`*_RELATIVE` relocations.
The `merge-dynamic.py` script has been improved to find the appropriate
spot to put the rewritten program/section and headers and the dynamic
sections (the kernel is peeky on them).
Finally the `setRegister` function has been introduced in the module
produced by `revamb`. This function allows to keep CSVs static and, at
the same time, it allow `support.c` to set them. This is particularly
useful when we want to call the `root` function with specific values in
the registers (e.g., during for fuzzing purposes) or, as it's the case
for PIE, to synchronize the value of the FS register, which is
initialized by the dynamic loader, before execution gets to the `main`
function in `support.c`.
The reaching definition analysis now considers loads and stores from and
to absolute addresses.
This improves the identification of target of indirect jumps in x86-64
PIC code.
This function computes the CSV that may be accessed from a call in root.
Before this commit, it assumed that all the load and store had to be
aligned with the underlying variables. This happens to be too strict an
assumption, and it is relaxed in this commit.
When a load or a store is not aligned with the underlying type, this
commit introduces code to mark all the spanned CSV as accessed.
This is information is then attached to the call site in root as
metadata, as it already happened before.
This commit improves the capabilities of the `CPUStateAccessAnalysis`.
In particular, it is capable of handling GEPs that access arrays at
unknown offsets. This is a common operation in QEMU helpers, because it
is often used to access some elements in the array of GPRs using indexes
that are not known at translate-time. Handling this case allows us to
shrink the size of the generated translated code, because we only
generate accesses at valid offsets in the arrays instead of handling all
the possible wild accesses in the CPU State.
This commit adds a new analysis pass: `CPUStateAccessAnalysisPass`.
This pass currently performs 4 operations.
1. A preliminary analysis of the call graph, to select the functions
that are reachable from the root function through direct calls. All
the other performed operations are executed on this set of reachable
functions.
2. An interprocedural forward taint analysis, starting from the uses of
`env`, the global variable pointing to the QEMU struct continaint the
CPU. This analysis taints all the instructions that use the address
of `env`, until a load or a store is met. If a load or a store uses a
tainted Value as address it means that it is accessing a CSV at a
given offset (which at this point is still unknown).
3. An interprocedural offset analysis, which deduces the possible
offsets used by every tainted load/store to access the CSV. This
analysis initially works backwards, exploring all the Values that
contribute at the computation of the addresses used by tainted
load/stores. Once it finds all the sources, it starts propagating the
values forward, collecting the offsets computed along the way. It
does this until it reaches the tainted load/stores again. At that
point the analysis knows all the possible offsets used by each
tainted load/store to access the CPU state.
4. The results of the previous steps are used to do 3 things:
* marking all the indirect calls with tainted arguments as illegal;
this is necessary because those calls may access the CPU State in
unpredictable ways;
* attaching metadata to all the call sites to QEMU helpers in the root
function; these metadata provide information on which parts of the
CPU State may be accessed from that call site, which is a
potentially useful information for users of libtinycode that we also
plan to use in other parts of revamb;
* substituting loads, stores, and memcpys to and from the CPU state
with accesses to global variables; this operation effectively
replaces what was previously done by the CorrectCPUStateUsagePass,
which is now obsolete and was removed in this commit.
This function takes a `Instruction *` pointing to a `CallInst`, and
returns a `Function *` to the Callee if it's a direct call, skipping all
the bitcasts if any.
If the argument is not a `CallInst`, or it's not a direct call, it
returns `nullptr`.
Check that the calling `unindent()` never reduces the indentation
"below" zero, causing `IndentLevel` to wrap around to high numbers.
If it drops below zero it's a bug anyway, so assert!
This commit does 6 things on `getTypeAtOffset()`:
1. it changes the second argument from `StructType *TheStruct` to a more
generic argument `Type *VarType`, making it capable of working on any
type;
2. it removes recursion, substituting it with a while loop;
3. it removes the now useless `Depth` argument;
4. it changes the return type to `std::pair<IntegerType *, unsigned>`,
because this was the assumption that all the callers did anyways;
5. it guards all the unexpected types with an assertion;
6. it purposely avoids to guard pointer types with assertions, as a
workaround for a specific situation documented in detail in the new
comments.
The function `getTypeSizeInBits()` was wrongly used in many places when
reasoning about memory allocation, memory accesses, and memory offsets,
The result was often divided by 8 (possibly losing spare bits) or even
multiplied by 8, which makes no sense.
These uses were error prone, even if they didn't cause problems yet.
The `getTypeAllocSize()` is better suited for these uses, because it
returns the number of bytes necessary to allocate an object of the given
Type.
This class introduces the `RunningStatistics` class, which allows to
compute the mean and standard deviation of a set of numbers. These
values are computed incrementally and can be associated to a name. The
values computed by `RunningStatistics` can be dumped upon regular
program termination, `SIGABRT` and `SIGINT`. In practice they are
printed at the end of the program execution, even in case of asserts and
`Ctrl + C`. Moreover, `SIGUSR1` is used to trigger printing the
statistics without crashing the program.
`Logger` is the new infrastructrure for debug messages. They are
supposed to be used by just creating a global variable in a translation
unit and then writing there directly as if they were a stream.
Key features:
* Self-registration: a `Logger` by default register itself automatically
in a register. This means that at run-time, unlike with `DBG`, we have
a list of all the possible `Logger`s. One benefit of this approach is
being able to list all the available `Logger`s in `--help`.
* Indentation: `Logger`s can be indented. In particular the
`LoggerIndent` helper class can automatically indent a certain
`Logger` during the lifetime of its instance.
* Atomic messages: a message is no longer simply delimited by a "\n": to
mark the end of a message, an instance of `LogTerminator` has to be
streamed to the `Logger`. This can also be associated directly to the
lifetime of an object by using the `LogOnReturn` class.
* Custom class formatting: it is possible to decide how a certain object
should be formatted when sent to a `Logger` by implementing the
`writeToLog` function. For example, `llvm::Value`-derived objects are
passed through the `getName` function. If no custom formatter is
specified, the `operator<<` is employed.
Sometimes, dealing with move-semantics in C++ can be challenging. To
mitigate this problem, this commit introduces `ClassSentinel`, a simple
class that can be added as a member of a class and that will allow the
user to easily monitor if an instance of a class is used after being
moved in an unwanted way. To do so, simply call the
`ClassSentinel::check` method on the class member.
`ClassSentinel` performs similar (but less reliable) checks on usage of
destroyed objects.
`ClassSentinel` can also (optionally, though the `SENTINEL_STACKTRACES`
macro) collect stack traces of the points where the object was moved or
destroyed.
This commit also includes some basic testing for the class.
`SmallMap` is a `std::map`-like data structure whose storage is inline
if the number of elements is small, pretty much as `llvm::SmallVector`
and friends. In case the `SmallMap` is in its small form, the search
cost is linear in the number of elements.
This commit also introduces `Iteratall`, an iterator class that can wrap
iterators of different types but with the same `::value_type` and
`::reference` types. This is used to abstract away the fact that
`SmallMap` iterators can either be iterators over a `std::array` or
`std::map`.
Note that if you want to iterate on the elements, and you want to be
sure that they are ordered (as happens with `std::map`), you have to
call the `sort` method before, which might or might not triggered a
sort. Note also that the internal inline storage uses a `std::array`,
which means that the `K` and `V` must be default constructible.
SET was predisposed to be able to track all the addresses from which a
load has been performed, but this has never been fully implemented. This
commit does that.
It is important to know that a load has been performed from a certain
address since, most likely, this means that part of the program is not
code. In fact, this information is used to tag jump targets, so that
other analysis can take into account this fact. Note however that, at
the current stage, the jump target itself is preserved (and, therefore,
translated).