Files
mandiant-GoReSym/objfile/strings_test.go
T
kamran ul haq d7f23b7a3e Add -strings flag to extract Go strings from binaries (#77)
* Add -strings flag and ExtractMetadata field

- Add Strings []string to ExtractMetadata struct
- Add -strings command-line flag for string extraction
- Update main_impl and main_impl_tmpfile signatures to accept printStrings parameter
- Add placeholder string extraction logic with TODO marker
- Update printForHuman to display extracted strings section
- Verified flag appears in help and outputs correctly in both JSON and human format

Part of #45

* Implement Go string extraction algorithm

- Create objfile/strings.go with core extraction logic
- Implement FLOSS-based string internment table detection
- Add string candidate scanning (pointer + length pairs)
- Implement findLongestMonotonicRun() for pattern detection
- Add UTF-8 validation and printability filtering
- Minimum string length: 4 characters, 80% printable
- Add helper methods to elfFile: getSections(), is64Bit(), isLittleEndian()
- Update main.go to call file.ExtractStrings() instead of placeholder
- Tested with testproject/testproject: extracts 512 strings successfully
- Extracts real Go strings: type names, runtime symbols, error messages

Based on FLOSS floss/language/go/extract.py algorithm
Part of #45

* Update README with -strings flag documentation

- Add -strings flag to available flags list
- Add Strings field to example JSON output
- Document purpose: extract embedded Go strings from binary

Part of #45

* Add tests validated against FLOSS output

Per maintainer request, added comprehensive test suite:

- strings_floss_test.go: Validates GoReSym against FLOSS reference output
  * 99.2% match rate (648/653 strings match FLOSS)
  * Uses FLOSS output from testproject.exe as ground truth
  * Reference saved in testdata/floss_reference.txt

- strings_test.go: Additional unit tests for:
  * ELF and PE binary string extraction
  * Monotonic run detection algorithm
  * String filtering (printability, minimum length)

- pe.go: Added helper methods (getSections, is64Bit, isLittleEndian)
  to enable string extraction from PE binaries

All tests pass.

* Address PR #77 review comments

- Convert getSections() to iterateSections() using callback pattern to avoid memory pressure
- Add Strings field to GoReSym.proto for external parsers
- Implement iterateSections() for Mach-O format (previously missing)

Changes requested by @stevemk14ebr in review:
1. Memory optimization: Replace array-based section loading with generator pattern
2. Proto definition: Add 'repeated string strings = 13' field
3. Mach-O support: Add missing iterateSections() implementation

* Fix test failures and expand string extraction testing per PR #77 feedback

- Fixed 4 test compilation errors by adding missing printStrings parameter to main_impl() calls
- Added comprehensive TestStringExtraction function with 7 test cases covering Linux/macOS/Windows binaries
- Implemented isPrintable() helper for ASCII validation (range 32-126)

* Align string extraction with FLOSS algorithm per PR #77 review

Rewrite string extraction to match FLOSS (extract.py) logic:
- Sort candidates by address (not length) to fix monotonic run detection
- Add image VA range and max section size filtering for candidates
- Use candidate (pointer, length) pairs for direct extraction from blob
- Replace 80% printable threshold with 100% fully printable check
- Fix PE section addresses to include ImageBase for correct VA comparison

Results: 648 strings extracted from PE test binary, 100% match with FLOSS.

* Run string extraction on test build files

* Extend main test to cover all new versions

* Fix string extraction for old Windows binaries and add length sanity check

1. Add .text to data sections for old Go Windows binaries (1.7-1.10)
   that store strings in the code section instead of .rdata

2. Add maxReasonableStringLength (64KB) to filter out garbage candidates
   with huge lengths that cause incorrect blob boundary detection

3. Update TestIsDataSection to reflect .text being a valid data section

Fixes string extraction failures on:
- Old Windows PE binaries (Go 1.7-1.10): 0 -> 256 strings
- Modern Windows: 574 strings
- macOS: 284 strings

All tests pass including TestWeirdBins on 15 real binaries.

* Fix test failures for Go 1.5/1.6 and Go 1.22 macOS; fix build script typo

1. main_test.go: Skip Strings check for Go 1.5/1.6
   Pre-SSA C-based linker does not produce length-sorted string blobs,
   so findLongestMonotonicRun never reaches the minimum threshold of 10.
   Pattern mirrors the existing interface-parsing guard.

2. objfile/strings.go: Handle prevNull == -1 in findStringBlobRange
   Apple's linker on newer Go/macOS (1.22+) packs sections without leading
   padding, so no null bytes exist before the first string candidate.
   bytes.LastIndex returns -1 in that case; treat it as offset 0 instead
   of bailing out with nil.

3. objfile/macho.go: Add missing is64Bit/isLittleEndian for machoFile
   elfFile and peFile already implement these interface methods; machoFile
   was the only rawFile implementation without them. Detect CPU type
   (CpuAmd64/Arm64 = 64-bit) and byte order from the macho.File struct.

4. build_test_files.sh: Fix $ver typo -> $GO_VER on mkdir line
   The directory was never created before Docker ran, causing Go 1.5
   builds to produce no output. Now all 12 binaries are built for
   every version including 1.5 and 1.6.

---------

Co-authored-by: Stephen Eckels <stevemk14ebr@gmail.com>
2026-02-17 10:14:48 -05:00

345 lines
10 KiB
Go

package objfile
import (
"path/filepath"
"testing"
)
// TestExtractStrings_ELF tests string extraction from an ELF binary (Linux)
func TestExtractStrings_ELF(t *testing.T) {
testBinary := filepath.Join("..", "testproject", "testproject")
file, err := Open(testBinary)
if err != nil {
t.Skipf("Could not open test binary: %v (run: cd testproject && go build)", err)
}
defer file.Close()
strings, err := file.ExtractStrings()
if err != nil {
t.Fatalf("ExtractStrings() failed: %v", err)
}
// Should extract a reasonable number of Go strings
if len(strings) < 100 {
t.Errorf("Expected at least 100 strings from ELF binary, got %d", len(strings))
}
// Verify common Go type/keyword strings are present
expectedStrings := []string{"func", "chan", "bool", "uint"}
assertStringsPresent(t, strings, expectedStrings)
}
// TestExtractStrings_PE tests string extraction from a PE binary (Windows)
func TestExtractStrings_PE(t *testing.T) {
testBinary := filepath.Join("..", "testproject", "testproject.exe")
file, err := Open(testBinary)
if err != nil {
t.Skipf("Could not open Windows test binary: %v (run: cd testproject && GOOS=windows GOARCH=amd64 go build -o testproject.exe)", err)
}
defer file.Close()
strings, err := file.ExtractStrings()
if err != nil {
t.Fatalf("ExtractStrings() failed: %v", err)
}
// Should extract a reasonable number of Go strings
if len(strings) < 100 {
t.Errorf("Expected at least 100 strings from PE binary, got %d", len(strings))
}
// Verify common Go type/keyword strings are present
expectedStrings := []string{"func", "chan", "bool", "uint"}
assertStringsPresent(t, strings, expectedStrings)
}
// TestFindLongestMonotonicRun validates the core algorithm
// Note: function requires at least 10 entries for a valid run (practical filter)
func TestFindLongestMonotonicRun(t *testing.T) {
t.Run("finds monotonic run in realistic data", func(t *testing.T) {
// Create realistic test data: monotonically increasing lengths
candidates := make([]StringCandidate, 50)
for i := range candidates {
candidates[i] = StringCandidate{
Pointer: uint64(0x1000 + i*10),
Length: uint64(i + 4), // lengths 4, 5, 6, ... 53
}
}
start, end := findLongestMonotonicRun(candidates)
// Should find the entire sequence as one run
if start == -1 || end == -1 {
t.Error("Expected to find a valid run, got (-1, -1)")
}
runLength := end - start + 1
if runLength < 10 {
t.Errorf("Expected run length >= 10, got %d", runLength)
}
})
t.Run("too short returns -1", func(t *testing.T) {
candidates := []StringCandidate{
{Length: 1}, {Length: 2}, {Length: 3}, // Only 3 entries
}
start, end := findLongestMonotonicRun(candidates)
if start != -1 || end != -1 {
t.Errorf("Expected (-1, -1) for short input, got (%d, %d)", start, end)
}
})
}
// TestIsFullyPrintable validates the printability filter.
// All characters must be printable (unicode.IsPrint) or common whitespace (\t, \n, \r).
func TestIsFullyPrintable(t *testing.T) {
tests := []struct {
input string
want bool
}{
{"hello world", true},
{"runtime.error", true},
{"line1\nline2", true}, // newlines are allowed
{"col1\tcol2", true}, // tabs are allowed
{"windows\r\n", true}, // carriage return allowed
{"\x00\x01\x02\x03", false}, // all non-printable
{"abc\x00\x01", false}, // any non-printable fails (was 80% threshold)
{"mostly ok \x01", false}, // even one non-printable byte fails
{"", false}, // empty string
}
for _, tt := range tests {
got := isFullyPrintable(tt.input)
if got != tt.want {
t.Errorf("isFullyPrintable(%q) = %v, want %v", tt.input, got, tt.want)
}
}
}
// TestExtractStrings_MinLength validates minimum length filtering
func TestExtractStrings_MinLength(t *testing.T) {
testBinary := filepath.Join("..", "testproject", "testproject")
file, err := Open(testBinary)
if err != nil {
t.Skip("Test binary not available")
}
defer file.Close()
strings, err := file.ExtractStrings()
if err != nil {
t.Fatalf("ExtractStrings() failed: %v", err)
}
// All strings must be >= 4 characters (MIN_STRING_LENGTH)
for _, s := range strings {
if len(s) < 4 {
t.Errorf("Found string shorter than 4 chars: %q (len=%d)", s, len(s))
}
}
}
// TestFindLongestMonotonicRun_MixedData tests that the algorithm correctly
// finds the blob candidates among noise. This simulates address-sorted
// candidates where only the middle portion points into the string blob.
func TestFindLongestMonotonicRun_MixedData(t *testing.T) {
// Simulate: 5 noise candidates, then 20 blob candidates, then 5 noise
var candidates []StringCandidate
// Noise before blob (random lengths, not monotonic)
noise1 := []uint64{50, 3, 100, 7, 42}
for i, l := range noise1 {
candidates = append(candidates, StringCandidate{
Pointer: uint64(0x1000 + i*8),
Length: l,
})
}
// String blob candidates: sorted by address, monotonically increasing lengths
// This is what Go's string internment table looks like
for i := 0; i < 20; i++ {
candidates = append(candidates, StringCandidate{
Pointer: uint64(0x4000 + i*10),
Length: uint64(4 + i), // 4, 5, 6, ..., 23
})
}
// Noise after blob
noise2 := []uint64{2, 80, 1, 60, 5}
for i, l := range noise2 {
candidates = append(candidates, StringCandidate{
Pointer: uint64(0x8000 + i*8),
Length: l,
})
}
start, end := findLongestMonotonicRun(candidates)
if start == -1 || end == -1 {
t.Fatal("Expected to find a valid run, got (-1, -1)")
}
runLength := end - start + 1
if runLength != 20 {
t.Errorf("Expected run length 20 (the blob candidates), got %d", runLength)
}
// The run should start at index 5 (after the 5 noise entries)
if start != 5 {
t.Errorf("Expected run to start at index 5, got %d", start)
}
}
// TestDeduplicateUint64 validates the deduplication helper
func TestDeduplicateUint64(t *testing.T) {
tests := []struct {
name string
input []uint64
expect []uint64
}{
{"no duplicates", []uint64{1, 2, 3}, []uint64{1, 2, 3}},
{"with duplicates", []uint64{1, 1, 2, 3, 3, 3, 4}, []uint64{1, 2, 3, 4}},
{"all same", []uint64{5, 5, 5}, []uint64{5}},
{"single element", []uint64{42}, []uint64{42}},
{"empty", []uint64{}, []uint64{}},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got := deduplicateUint64(tt.input)
if len(got) != len(tt.expect) {
t.Errorf("deduplicateUint64(%v) returned %d elements, want %d", tt.input, len(got), len(tt.expect))
return
}
for i := range got {
if got[i] != tt.expect[i] {
t.Errorf("deduplicateUint64(%v)[%d] = %d, want %d", tt.input, i, got[i], tt.expect[i])
}
}
})
}
}
// TestIsDataSection validates section name matching across ELF/PE/Mach-O formats
func TestIsDataSection(t *testing.T) {
tests := []struct {
name string
want bool
}{
// ELF sections
{".rodata", true},
{".data", true},
{".noptrdata", true},
// Mach-O sections
{"__rodata", true},
{"__data", true},
{"__noptrdata", true},
// PE sections
{".rdata", true},
{".text", true}, // Included for old Go Windows binaries (1.7-1.10) that store strings in .text
// Suffixed variants (some formats append .__)
{".rodata.__", true},
// Non-data sections
{".bss", false},
{"__TEXT", false},
{".got", false},
}
for _, tt := range tests {
got := isDataSection(tt.name)
if got != tt.want {
t.Errorf("isDataSection(%q) = %v, want %v", tt.name, got, tt.want)
}
}
}
// TestFindStringCandidates validates candidate extraction from raw binary data
func TestFindStringCandidates(t *testing.T) {
t.Run("64-bit little-endian", func(t *testing.T) {
// Create a fake section with one string struct:
// struct: pointer=0x5000, length=5
// Scanner steps at ptrSize (8-byte) alignment, so we use exactly
// one 16-byte struct to avoid overlapping false positives.
data := make([]byte, 16)
// pointer = 0x5000 (little-endian uint64)
data[0] = 0x00
data[1] = 0x50
data[2] = 0x00
data[3] = 0x00
data[4] = 0x00
data[5] = 0x00
data[6] = 0x00
data[7] = 0x00
// length = 5 (little-endian uint64)
data[8] = 0x05
data[9] = 0x00
data[10] = 0x00
data[11] = 0x00
data[12] = 0x00
data[13] = 0x00
data[14] = 0x00
data[15] = 0x00
// imageMin=0x1000, imageMax=0x10000 (pointer 0x5000 is in range), maxSectionSize=1000
candidates := findStringCandidates(data, 0x1000, true, true, 0x1000, 0x10000, 1000)
if len(candidates) != 1 {
t.Fatalf("Expected 1 candidate, got %d", len(candidates))
}
if candidates[0].Pointer != 0x5000 || candidates[0].Length != 5 {
t.Errorf("candidate[0] = {%#x, %d}, want {0x5000, 5}", candidates[0].Pointer, candidates[0].Length)
}
})
t.Run("skips zero pointer and length", func(t *testing.T) {
data := make([]byte, 16)
// pointer=0, length=0 -- should be skipped
candidates := findStringCandidates(data, 0x1000, true, true, 0x1000, 0x10000, 1000)
if len(candidates) != 0 {
t.Errorf("Expected 0 candidates for zero data, got %d", len(candidates))
}
})
t.Run("skips pointer outside image range", func(t *testing.T) {
data := make([]byte, 16)
// pointer = 0x5000 (outside range [0x1000, 0x4000))
data[0] = 0x00
data[1] = 0x50
data[8] = 0x05
candidates := findStringCandidates(data, 0x1000, true, true, 0x1000, 0x4000, 1000)
if len(candidates) != 0 {
t.Errorf("Expected 0 candidates for out-of-range pointer, got %d", len(candidates))
}
})
t.Run("skips length exceeding max section size", func(t *testing.T) {
data := make([]byte, 16)
// pointer = 0x2000 (in range), length = 500 (exceeds maxSectionSize=100)
data[0] = 0x00
data[1] = 0x20
data[8] = 0xF4 // 500 in little-endian
data[9] = 0x01
candidates := findStringCandidates(data, 0x1000, true, true, 0x1000, 0x10000, 100)
if len(candidates) != 0 {
t.Errorf("Expected 0 candidates for oversized length, got %d", len(candidates))
}
})
}
// Helper function to check if expected strings are present
func assertStringsPresent(t *testing.T, strings []string, expected []string) {
t.Helper()
stringSet := make(map[string]bool)
for _, s := range strings {
stringSet[s] = true
}
for _, exp := range expected {
if !stringSet[exp] {
t.Errorf("Expected string %q not found in extracted strings", exp)
}
}
}