BinTopsy (Binary Autopsy) is a collection of lightweight Python scripts that assist in the early stages of static malware analysis and reverse engineering.
The toolkit lets you visualise file entropy, disassemble code snippets on the fly, scan binaries with YARA rules (paginated for large dumps), automate threat-intelligence lookups via VirusTotal, and dissect binaries with radare2 — extracting structural data, generating control-flow / call graphs, tracing API usage, and clustering functions via fuzzy hashes.
| # | Script | Purpose | Key deps |
|---|---|---|---|
| 1 | entropy-viz.py |
2D Shannon-entropy heatmap of a binary | numpy, matplotlib |
| 2 | disasm.py |
CLI disassembler (x86/x64/ARM/MIPS) for files & hex strings | capstone |
| 3 | yara-chunk-scanner.py |
Chunked, recursive YARA scanner; parallel directory mode | yara-python |
| 4 | vt-folder-scan.py |
Hash-only VirusTotal triage of a folder, with local cache | requests, dotenv |
| 5 | r2-dissector.py |
Radare2 dissector → JSON + CFG / call-graph PDFs + HTML index | r2pipe, graphviz |
| 6 | sda-hashes.py |
Enriches dissector JSON with TLSH / ssdeep fuzzy hashes | python-tlsh, ssdeep |
| 7 | r2-call-tracer.py |
Brute-force call tracer (also resolves indirect IAT calls) | r2pipe |
| 8 | json-behavior-analyzer.py |
Behavioural matrix + MITRE ATT&CK Navigator layer | stdlib |
| 9 | r2-call-graph.py |
Targeted call graph rooted at a specific function (BFS depth) | r2pipe, graphviz |
| 10 | r2-xref-grapher.py |
CFGs of every function calling a given API | r2pipe, graphviz |
| 11 | bintopsy-cluster.py |
Cluster similar functions across enriched JSONs | python-tlsh / ssdeep |
| 12 | bintopsy-diff.py |
Structural diff between two dissector JSONs (with rename det.) | python-tlsh (opt.) |
| 13 | bintopsy-report.py |
Single-command pipeline → self-contained HTML report | all of the above |
Some Python packages wrap native libraries that must be present before
pip install:
| Tool / Library | Required by | Debian / Ubuntu | macOS (Homebrew) |
|---|---|---|---|
| Radare2 | r2-dissector.py, r2-call-*.py, r2-xref-grapher.py |
apt install radare2 |
brew install radare2 |
| Graphviz | r2-dissector.py, r2-call-graph.py, r2-xref-grapher.py |
apt install graphviz |
brew install graphviz |
| libfuzzy | sda-hashes.py (ssdeep) |
apt install libfuzzy-dev |
brew install ssdeep |
git clone https://github.com/reverseame/BinTopsy.git
cd BinTopsy
pip install -r requirements.txtOnly needed if you intend to use vt-folder-scan.py.
Create a file named secrets.env in the repo root (already gitignored):
VT_API_KEY=your_64_character_api_key_hereA. Quick overview of a packed executable — default 4 KB window, default 64-column grid:
python entropy-viz.py samples/malware.exe -o overview.pdfB. High-resolution analysis (shellcode / steganography) — 256-byte window plus a wider grid for a denser map:
python entropy-viz.py suspicious_image.png -w 256 -W 128 -o high_res_map.pdfPadding at the end of the file is now masked (transparent), so it no longer looks like a fake low-entropy region.
A. Shellcode from a hex dump — disassemble inline, no file needed:
python disasm.py -s "55 48 89 e5 48 83 ec 20" -a x64B. IoT firmware (MIPS, big-endian) — set base address and endianness:
python disasm.py -f firmware_bootloader.bin -a mips --big-endian --base 0x80001000C. Limit output — first 30 instructions only:
python disasm.py -f sample.bin -a x86 -n 30A. Memory dump — scan an 8 GB raw dump in 4 MB chunks:
python yara-chunk-scanner.py memory_dump.raw ./rules/malware.yar -p 4194304B. Bulk directory scan — recursive directory of files vs. directory of rules, in parallel across 8 workers:
python yara-chunk-scanner.py ./extracted_files/ ./rules_repo/ -p 4096 -j 8C. Disable cross-chunk overlap — for performance when rules have only short strings:
python yara-chunk-scanner.py ./big_blob.bin ./rules.yar -p 65536 -O 0The default 1 KB overlap catches strings that straddle the boundary between two chunks; the previous version silently missed them.
A. Triage a Downloads folder, ignoring common media files:
python vt-folder-scan.py ~/Downloads --avoid .jpg .jpeg .png .mp4 .log .txtB. Premium key + JSON report:
python vt-folder-scan.py ./incident_response_data --premium -o triage.jsonC. Limit by file size to save quota:
python vt-folder-scan.py ./suspicious --max-size 10485760 # skip > 10 MBD. Cache — results are persisted to .vt_cache.json (TTL 7 days by
default), so re-running on the same folder is free. Tweak with
--cache-ttl 30 or disable with --no-cache.
A. Full JSON export — metadata, entropy and basic blocks per function:
python r2-dissector.py samples/malware.exe --blocks > structural_analysis.jsonB. Visualization (call graph) — global map of who-calls-whom:
python r2-dissector.py samples/malware.exe --call-graph -od ./graphsC. CFG of a single function:
python r2-dissector.py samples/malware.exe -g sym.main -od ./graphsAdds TLSH and ssdeep hashes to the JSON produced by r2-dissector.py:
python sda-hashes.py structural_analysis.json
# → writes structural_analysis_enriched.jsonUse only one algorithm:
python sda-hashes.py structural_analysis.json --tlshWalks every function and records direct + indirect (IAT) calls:
python r2-call-tracer.py samples/obfuscated.bin > call_trace.jsonCompact JSON (smaller, no indentation):
python r2-call-tracer.py samples/obfuscated.bin --compact > call_trace.jsonA. Triage — only show functions with interesting capabilities:
python json-behavior-analyzer.py call_trace.jsonOutput highlights (colours):
- Yellow → file + network (downloader / dropper)
- Red → file + crypto (ransomware-like)
- Magenta → process injection (
VirtualAllocEx,WriteProcessMemory, etc.) - Cyan → heavy anti-analysis (≥3 such APIs)
B. Verbose audit — show every function, including GUI-only / generic ones:
python json-behavior-analyzer.py call_trace.json --verboseC. ATT&CK Navigator layer + machine-readable findings:
python json-behavior-analyzer.py call_trace.json \
--json findings.json --attack-layer attack_layer.jsonDrag attack_layer.json into https://mitre-attack.github.io/attack-navigator/
to visualise techniques observed in the sample.
POSIX equivalents (
open,socket,ptrace,mmap,dlopen, …) are mapped alongside the Windows API names, so the analyser is useful for Linux and macOS samples too.
Generates a focused call graph rooted at a specific function. Complements
r2-dissector.py --call-graph, which produces the global graph.
A. From main, depth 4:
python r2-call-graph.py samples/malware.exe main_subgraph.pdf -s main -d 4B. From an address:
python r2-call-graph.py samples/malware.exe entry.pdf -s 0x401000 -d 5Nodes are coloured by cyclomatic complexity (CC): green = root, blue ≤ 10, yellow > 10, orange > 20, red > 50.
A. List every import with its xref count and the maximum CC of its callers (rapid triage of "where is the danger?"):
python r2-xref-grapher.py samples/malware.exe -lB. Generate a CFG per caller of a specific API:
python r2-xref-grapher.py samples/malware.exe VirtualAllocEx -o ./xref_graphsC. Substring matching (e.g. catch both RegOpenKey and RegOpenKeyEx):
python r2-xref-grapher.py samples/malware.exe RegOpenKey --partialDefault matching is exact (case-insensitive) to avoid false positives.
Group near-duplicate functions across one or more enriched JSONs (output of
sda-hashes.py). Useful for finding shared code between malware variants
or detecting statically-linked libraries.
A. Single-binary cluster — find duplicate functions inside one sample:
python bintopsy-cluster.py samples/malware_enriched.json -t 60B. Cross-sample family detection — feed several enriched JSONs:
python bintopsy-cluster.py family/*_enriched.json -t 70 \
--csv shared_code.csv --dendrogram tree.pdfCross-sample clusters are tagged (cross-sample), surfacing the most
interesting matches first. The --dendrogram flag is optional and needs
SciPy + Matplotlib.
Compare two dissector JSONs side by side: added / removed / modified functions. If both inputs were enriched with TLSH, the script will also flag likely renames (a removal-and-addition pair whose TLSH distance is low enough to be the same code under a different name).
python bintopsy-diff.py samples/v1_enriched.json samples/v2_enriched.jsonSave the diff as JSON for downstream tooling:
python bintopsy-diff.py v1.json v2.json --json diff.jsonOne command runs the full BinTopsy stack on a binary and produces a self-contained HTML report indexing every artifact:
python bintopsy-report.py samples/malware.exe -o ./report
open ./report/index.htmlThe pipeline runs: dissector → fuzzy hashing → call tracer → behavioural matrix + ATT&CK layer → entropy heatmap → CFG gallery → global call graph.
Useful flags:
--no-graph-allskip the per-function CFG batch (slow on big binaries).--no-call-graphskip the global call graph.--yara rules.yaradd a YARA scan to the pipeline.
pip install pytest
pytest tests/ -vThe suite is split in two:
tests/test_smoke.py— every script must run--helpcleanly.tests/test_functional.py— compiles a tiny C program once per session (skipped if no compiler is onPATH) and runs each tool against it, verifying outputs are well-formed.
Tests that depend on radare2, dot or yara skip cleanly when those
binaries are missing.
The radare2 toolchain is designed to compose. A typical analysis flow is:
┌─────────────────────┐
binary ──────────► │ r2-dissector.py │ ──► structural_analysis.json
└─────────────────────┘ │
▼
┌─────────────────────┐
│ sda-hashes.py │ ──► *_enriched.json
└─────────────────────┘ (TLSH / ssdeep
per function)
┌─────────────────────┐
binary ──────────► │ r2-call-tracer.py │ ──► call_trace.json
└─────────────────────┘ │
▼
┌────────────────────────────┐
│ json-behavior-analyzer.py │ ──► capability matrix
└────────────────────────────┘
┌─────────────────────┐
binary ──────────► │ r2-xref-grapher.py │ ──► per-API caller CFGs (PDF)
└─────────────────────┘
┌─────────────────────┐
│ r2-call-graph.py │ ──► targeted call graph (PDF)
└─────────────────────┘
For one-shot analysis of a single binary, bintopsy-report.py orchestrates
the whole tree and emits a self-contained HTML report:
binary ──► bintopsy-report.py ──► report/index.html
├── dissect.json
├── dissect_enriched.json
├── calls.json
├── behaviour.json
├── attack_layer.json ← MITRE Navigator
├── entropy.pdf
└── graphs/ ← CFGs + call graph
└── index.html
For comparing or clustering several samples:
*_enriched.json (N files) ──► bintopsy-cluster.py ──► clusters / dendrogram
v1.json + v2.json ──► bintopsy-diff.py ──► added/removed/renamed
The methodology behind BinTopsy is detailed in:
Toward Structured Memory Forensics: A MITRE ATT&CK-Aligned Workflow for Malware Investigation (Ricardo J. Rodríguez) Published in: 16th International Conference on Digital Forensics & Cyber Crime (ICDF2C 2025) Web repository
@InProceedings{Rodriguez-ICDF2C-25,
author = {Ricardo J. Rodríguez},
booktitle = {Proceedings of the 16th EAI International Conference on Digital Forensics & Cyber Crime},
title = {{Toward Structured Memory Forensics: A MITRE ATT\&CK-Aligned Workflow for Malware Investigation}},
year = {2025},
note = {Accepted for publication. To appear.},
number = {PP},
pages = {PP},
publisher = {Springer},
volume = {PP},
abstract = {Memory forensics is emerging as an essential technique for detecting malware-related volatile indicators of compromise (IoCs) that traditional disk analysis may miss. However, the lack of standardized best practices for analyzing memory-resident malware evidence continues to limit the effectiveness and reproducibility of forensic investigations. In this work, we propose a structured five-phase workflow that formalizes best practices for the extraction and analysis of malware-related IoCs, from initial evidence preservation to binary program investigation. Our methodology is explicitly aligned with the MITRE ATT\&CK framework, allowing analysts to correlate volatile memory artifacts with known adversarial tactics and techniques. Additionally, we examine technical challenges (such as paging, on-demand paging, memory inconsistencies, and runtime binary transformations) that threaten the integrity and reliability of memory evidence. We further propose practical recommendations and outline future research directions for addressing these challenges, with the goal of improving the reliability, consistency, and forensic robustness of memory-based malware analysis.},
keywords = {digital forensics, memory forensics, methodology, indicators of compromise, malware analysis},
url = {https://webdiis.unizar.es/~ricardo/files/papers/Rodriguez-ICDF2C-25.pdf},
}This repository is an example of human-AI collaboration.
- Code: the Python scripts were drafted by Google Gemini, then rigorously reviewed, refactored and tested by R. J. Rodríguez to ensure accuracy and safety.
- Visual assets: the BinTopsy logo was designed and generated with Google Gemini's image-generation capabilities.
- Architecture & methodology: the tool selection, scanning logic and research methodology were defined by R. J. Rodríguez.
Disclaimer: while AI was used to accelerate the development process, every line of code has been manually audited by the author. The resulting tools represent a verified implementation of the concepts described in the methodology.
Part of this research was supported by the Spanish National Cybersecurity Institute (INCIBE) under Proyecto Estratégico de Ciberseguridad — CIBERSEGURIDAD EINA UNIZAR and by the Recovery, Transformation and Resilience Plan funds, financed by the European Union (Next Generation).

