Files
lifting-bits-remill/test_runner_lib/TestRunner.cpp
T
Alex Cameron fb018c96e9 PowerPC Support (#645)
* Add skeleton for PPC

* Copyright notices

* Fill in some details for the PPC arch

* Start building a (wrong) PPC runtime

* Begin populating state structure

* First pass for EIS state structure

* Map registers to Sleigh register names

* More fixes

* add optional param

* Create handle unsupported and invalid instruction isels

* Correct typo

* Get a basic `remill-lift` invocation running without failure

* Fix capitalisation

* Set vle context reg

* Fix SleighDecoder signatures

* Set VLE context register in the Sleigh engine in addition to our
internal context reg mapping

* Capitalize reg names

* Add the flag registers for XER and CR

* Rename bitflag structures in PPC state

* PPC Sleigh patches (#643)

* Modified sleigh patch script to generate patches for multiple .sinc files

* update README with new examples of sleigh patch script invocation

* add ppc register definition

* add ppc sleigh patches

* fix issue with remill_insn_size definition

* regenerate sleigh patches for PPC

* update CMakeLists.txt to include PPC patches

* Add TEA signal as a register in the PPC state

* Uppercase the stack pointer register name

* Fix PPC instruction sizes

* initial PPC tests

* remove duplicate tests

* fix tests for e_stmvgprw/e_ldmvgprw

* add tests for loading/storing from special registers

* add tests with internal conditionals in pcode

* fix for pc reg and addr width not being the same... I suspect this issue is going to come up elsewhere

* add heuristic for flow from normal intrainstruction flow

* rework tests to allow testing for different sized registers

* add tests for overflow and record add

* fix bug with log printout

* add intrafunction control flow lifting

* handle edge case where there is no pcode op at the zero index

* Fix another inconsistency with mismatching address and PC reg size

* Allocate unique ptrs in the entry block

* Fix `INT_LEFT` and `INT_RIGHT` impl where shift exceeds bit width

* fix supiece lift?

* Add PPC emulate instruction to hyper call

* fix for pc reg and addr width not being the same... I suspect this issue is going to come up elsewhere

* add heuristic for flow from normal intrainstruction flow

* add intrafunction control flow lifting

* handle edge case where there is no pcode op at the zero index

* Fix another inconsistency with mismatching address and PC reg size

* fix supiece lift?

* Allocate unique ptrs in the entry block

* Fix `INT_LEFT` and `INT_RIGHT` impl where shift exceeds bit width

* fix int2float semantics

should use appropriate sized float based on the output size

* add tests for lifting int2float

* fix INT_{LEFT,RIGHT} semantics

should be `ICmpSGE` instead of `ICmpSGT`

* add cr0-7 registers

* fix formatting

* fix conditional branch test

* add test for compare

* re-enable rotate left word immediate and mask test

* genericize TestSpecOutput

* explicit instruction data size

* add test for syscall/callother (disabled)

* add tests for store/load word

* add test to convert from float to int

* specify intrinsic arg type, fixes null deref

* Add PPC emulate instruction to hyper call

* add headers + formatting

* remove old comment

* Map CRALL register

* Add basic LLVM data layout that specifies 32-bit addresses

* Remove unused variables

* convert auto* to auto when possible

* RegisterPrecondition -> RegisterCondition

* fix variable name

* convert any to variant

* use std::move

* bump to c++20, use concepts

* set arch in constructor since class isn't generic anyways

* formatting

* make type aliases

* bump cxx-common

* add comment

* clang format

* throw exception if register not found

* use const ref

* use shorthand for lambda capture values

* Add more detail to data layout to include proper stack alignment

* Compare to the correct size for SUBPIECE impl

* Add Sleigh message to error

* throw exception in else case

* throw runtime error if register value has incorrect type

* use reference instead of value

* get rid of unnecessary type alias

* formatting

* Propagate VLE context reg value into Sleigh

* Remove unnecessary whitespace

* Remove stale TODO and NOTE comments

* add additional parameter to test runner to specify decoding context

* drop llvm 14, bump macos version

* bump cxx-common, fix ci.yml mac build

* add test for unconditional relative negative branch

* add missing space to pcode debug log

* fix bug due to unordered_map, iteration order matters

* add error log in case we aren't able to adjust PC value

* use helper for getting register reference

* Revert "add optional param"

This reverts commit 51ed49f8cf.

* Remove remaining LLVM 14 compatibility code and configuration

* Add padding between CR and XER flags

* Use `enum class`

* Remove void cast

* Remove unnecessary variable

* Use initialiser lists where appropriate

* Remove redundant `else`

* Prefer `CHECK` over `assert`

* Polish PowerPC function initialisation with lambda

* zero out xer_so to fix tests

* log error when we see claim_eq with no usages

* Collapse namespace blocks

* Remove unnecessary `this->`

* Use `auto` where appropriate

* Remove unnecessary `else`

* Use `emplace` over `insert` for `std::map`

Co-authored-by: lkorenc <lukas.korencik@trailofbits.com>

* Use `constexpr` for VLE reg name

* Use module verification util

* Add `VerifyFunction` util and use where applicable

* Extract lambda to improve readability of flow categorisation

* Use type alias for context values

* Introduce type alias for block exit

* Create type alias for optional branch taken

* Refactor `PcodeCFGBuilder`

* Use lambda to avoid conditional mutation

* Extract duplicated bit-shift code generation into helper

* Simplify flow with ternary

* Add `GetBlock` helper

* Rename variables

* Move statement for clarity

* Create helpers for working with Sleigh context register values

* Convert loop to `std::copy`

* Add a comment explaining the use of set to de-duplicate and sort

* Refactor `IntraProcTransferCollector`

* Expose static method to easily use `IntraProcTransferCollector`

* Rename PPC related variables to include address width

* add docs to intrainstructionindex

* remove llvm 14 ifdefs

* don't log error if no claim_eqs were used

* update comments

* Cleanup exit visitors

---------

Co-authored-by: 2over12 <ian.smith@trailofbits.com>
Co-authored-by: William Tan <1284324+Ninja3047@users.noreply.github.com>
Co-authored-by: lkorenc <lukas.korencik@trailofbits.com>
2023-02-02 10:05:28 -05:00

315 lines
10 KiB
C++

/*
* Copyright (c) 2022 Trail of Bits, Inc.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#include <glog/logging.h>
#include <gtest/gtest.h>
#include <llvm/ExecutionEngine/ExecutionEngine.h>
#include <llvm/ExecutionEngine/GenericValue.h>
#include <llvm/ExecutionEngine/Interpreter.h>
#include <llvm/ExecutionEngine/MCJIT.h>
#include <llvm/IR/Instructions.h>
#include <llvm/IR/Verifier.h>
#include <llvm/Passes/PassBuilder.h>
#include <llvm/Support/DynamicLibrary.h>
#include <llvm/Support/Endian.h>
#include <llvm/Support/FileSystem.h>
#include <llvm/Support/JSON.h>
#include <llvm/Support/TargetSelect.h>
#include <llvm/Transforms/Utils/Cloning.h>
#include <remill/Arch/Arch.h>
#include <remill/BC/InstructionLifter.h>
#include <remill/BC/IntrinsicTable.h>
#include <remill/BC/Lifter.h>
#include <remill/BC/SleighLifter.h>
#include <remill/BC/Util.h>
#include <test_runner/TestRunner.h>
#include <random>
namespace test_runner {
namespace {
static bool FuncIsIntrinsicPrefixedBy(const llvm::Function *func,
const char *prefix) {
return func->isDeclaration() && func->getName().startswith(prefix);
}
} // namespace
uint8_t random_boolean_flag(random_bytes_engine &rbe) {
std::uniform_int_distribution<> gen(0, 1);
return gen(rbe);
}
void *MissingFunctionStub(const std::string &name) {
if (auto res = llvm::sys::DynamicLibrary::SearchForAddressOfSymbol(name)) {
return res;
}
#ifdef __APPLE__
if (name.at(0) == '_') {
if (auto res = llvm::sys::DynamicLibrary::SearchForAddressOfSymbol(
name.substr(1, name.length()))) {
return res;
}
}
#endif
LOG(FATAL) << "Missing function: " << name;
return nullptr;
}
MemoryHandler::MemoryHandler(llvm::support::endianness endian_)
: endian(endian_) {}
MemoryHandler::MemoryHandler(
llvm::support::endianness endian_,
std::unordered_map<uint64_t, uint8_t> initial_state)
: state(std::move(initial_state)),
endian(endian_) {}
uint8_t MemoryHandler::read_byte(uint64_t addr) {
if (state.find(addr) != state.end()) {
return state.find(addr)->second;
}
auto genned = rbe();
uninitialized_reads.insert({addr, genned});
state.insert({addr, genned});
return genned;
}
std::vector<uint8_t> MemoryHandler::readSize(uint64_t addr, size_t num) {
std::vector<uint8_t> bytes;
for (size_t i = 0; i < num; i++) {
bytes.push_back(this->read_byte(addr + i));
}
return bytes;
}
const std::unordered_map<uint64_t, uint8_t> &MemoryHandler::GetMemory() const {
return this->state;
}
std::string MemoryHandler::DumpState() const {
llvm::json::Object mapping;
for (const auto &kv : this->state) {
std::stringstream ss;
ss << kv.first;
mapping[ss.str()] = kv.second;
}
std::string res;
llvm::json::Value v(std::move(mapping));
llvm::raw_string_ostream ss(res);
ss << v;
return ss.str();
}
std::unordered_map<uint64_t, uint8_t> MemoryHandler::GetUninitializedReads() {
return this->uninitialized_reads;
}
extern "C" {
uint8_t __remill_undefined_8(void) {
return 0;
}
uint8_t __remill_read_memory_8(MemoryHandler *memory, uint64_t addr) {
LOG(INFO) << "Reading " << std::hex << addr;
auto res = memory->ReadMemory<uint8_t>(addr);
LOG(INFO) << "Read memory " << res;
return res;
}
MemoryHandler *__remill_write_memory_8(MemoryHandler *memory, uint64_t addr,
uint8_t value) {
LOG(INFO) << "Writing " << std::hex << addr
<< " value: " << (unsigned int) value;
memory->WriteMemory<uint8_t>(addr, value);
return memory;
}
uint16_t __remill_read_memory_16(MemoryHandler *memory, uint64_t addr) {
LOG(INFO) << "Reading " << std::hex << addr;
auto res = memory->ReadMemory<uint16_t>(addr);
LOG(INFO) << "Read memory " << res;
return res;
}
MemoryHandler *__remill_write_memory_16(MemoryHandler *memory, uint64_t addr,
uint16_t value) {
LOG(INFO) << "Writing " << std::hex << addr << " value: " << value;
memory->WriteMemory<uint16_t>(addr, value);
return memory;
}
uint32_t __remill_read_memory_32(MemoryHandler *memory, uint64_t addr) {
LOG(INFO) << "Reading " << std::hex << addr;
auto res = memory->ReadMemory<uint32_t>(addr);
LOG(INFO) << "Read memory " << std::hex << res;
return res;
}
MemoryHandler *__remill_write_memory_32(MemoryHandler *memory, uint64_t addr,
uint32_t value) {
LOG(INFO) << "Writing " << std::hex << addr << " value: " << value;
memory->WriteMemory<uint32_t>(addr, value);
return memory;
}
uint64_t __remill_read_memory_64(MemoryHandler *memory, uint64_t addr) {
LOG(INFO) << "Reading " << std::hex << addr;
return memory->ReadMemory<uint64_t>(addr);
}
MemoryHandler *__remill_write_memory_64(MemoryHandler *memory, uint64_t addr,
uint64_t value) {
LOG(INFO) << "Writing " << std::hex << addr << " value: " << value;
memory->WriteMemory<uint64_t>(addr, value);
return memory;
}
}
LiftingTester::LiftingTester(std::shared_ptr<llvm::Module> semantics_module_,
remill::OSName os_name, remill::ArchName arch_name)
: semantics_module(std::move(semantics_module_)),
arch(remill::Arch::Build(&this->semantics_module->getContext(), os_name,
arch_name)),
table(std::make_unique<remill::IntrinsicTable>(
this->semantics_module.get())) {
this->arch->InitFromSemanticsModule(this->semantics_module.get());
this->lifter = this->arch->DefaultLifter(*this->table.get());
}
LiftingTester::LiftingTester(llvm::LLVMContext &context, remill::OSName os_name,
remill::ArchName arch_name)
: arch(remill::Arch::Build(&context, os_name, arch_name)) {
this->semantics_module =
std::shared_ptr(remill::LoadArchSemantics(this->arch.get()));
this->table =
std::make_unique<remill::IntrinsicTable>(this->semantics_module.get());
this->lifter = this->arch->DefaultLifter(*this->table.get());
}
std::unordered_map<TypeId, llvm::Type *> LiftingTester::GetTypeMapping() {
std::unordered_map<TypeId, llvm::Type *> res;
res.emplace(TypeId::MEMORY, this->arch->MemoryPointerType());
res.emplace(TypeId::STATE, this->arch->StateStructType());
return res;
}
std::optional<std::pair<llvm::Function *, remill::Instruction>>
LiftingTester::LiftInstructionFunction(std::string_view fname,
std::string_view bytes, uint64_t address,
const remill::DecodingContext &ctx) {
remill::Instruction insn;
// This works for now since each arch has an initial context that represents the arch correctly.
if (!this->arch->DecodeInstruction(address, bytes, insn, ctx)) {
LOG(ERROR) << "Failed decode";
return std::nullopt;
}
LOG(INFO) << "Decoded insn " << insn.Serialize();
auto target_func =
this->arch->DefineLiftedFunction(fname, this->semantics_module.get());
LOG(INFO) << "Func sig: "
<< remill::LLVMThingToString(target_func->getType());
if (remill::LiftStatus::kLiftedInstruction !=
insn.GetLifter()->LiftIntoBlock(insn, &target_func->getEntryBlock())) {
target_func->eraseFromParent();
return std::nullopt;
}
auto mem_ptr_ref =
remill::LoadMemoryPointerRef(&target_func->getEntryBlock());
llvm::IRBuilder bldr(&target_func->getEntryBlock());
auto pc_ref = remill::LoadProgramCounterRef(&target_func->getEntryBlock());
auto next_pc_ref =
remill::LoadNextProgramCounterRef(&target_func->getEntryBlock());
bldr.CreateStore(
bldr.CreateLoad(llvm::IntegerType::get(target_func->getContext(), 32),
next_pc_ref),
pc_ref);
bldr.CreateRet(bldr.CreateLoad(this->lifter->GetMemoryType(), mem_ptr_ref));
return std::make_pair(target_func, insn);
}
std::optional<std::pair<llvm::Function *, remill::Instruction>>
LiftingTester::LiftInstructionFunction(std::string_view fname,
std::string_view bytes,
uint64_t address) {
return LiftInstructionFunction(fname, bytes, address,
this->arch->CreateInitialContext());
}
const remill::Arch::ArchPtr &LiftingTester::GetArch() const {
return this->arch;
}
static constexpr const char *kFlagIntrinsicPrefix = "__remill_flag_computation";
static constexpr const char *kCompareFlagIntrinsicPrefix = "__remill_compare";
/// NOTE(Ian): This stub is variadic to handle flag computations which accept arbitrary operand width types. since this function
/// stubs out to an identity function at runtime this is fine.
bool flag_computation_stub(bool res, ...) {
return res;
}
bool compare_instrinsic_stub(bool res) {
return res;
}
llvm::Function *
CopyFunctionIntoNewModule(llvm::Module *target, const llvm::Function *old_func,
const std::unique_ptr<llvm::Module> &old_module) {
auto new_f = llvm::Function::Create(old_func->getFunctionType(),
old_func->getLinkage(),
old_func->getName(), target);
remill::CloneFunctionInto(old_module->getFunction(old_func->getName()),
new_f);
return new_f;
}
void StubOutFlagComputationInstrinsics(llvm::Module *mod,
llvm::ExecutionEngine &exec_engine) {
for (auto &func : mod->getFunctionList()) {
if (FuncIsIntrinsicPrefixedBy(&func, kFlagIntrinsicPrefix)) {
exec_engine.addGlobalMapping(&func, (void *) &flag_computation_stub);
}
if (FuncIsIntrinsicPrefixedBy(&func, kCompareFlagIntrinsicPrefix)) {
exec_engine.addGlobalMapping(&func, (void *) &compare_instrinsic_stub);
}
}
}
} // namespace test_runner