Skip to content

Commit 0857ad1

Browse files
authored
Merge pull request #380 from Cisco-Talos/dev
chore: release EvidenceForge 1.15.0
2 parents eeb6225 + 41a1412 commit 0857ad1

3,743 files changed

Lines changed: 20135252 additions & 5564 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

.github/workflows/ci.yml

Lines changed: 4 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -24,7 +24,7 @@ jobs:
2424
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
2525

2626
- name: Install uv
27-
uses: astral-sh/setup-uv@37802adc94f370d6bfd71619e3f0bf239e1f3b78 # v7
27+
uses: astral-sh/setup-uv@c771a70e6277c0a99b617c7a806ffedaca235ff9 # v9.0.0
2828

2929
- name: Set up Python
3030
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
@@ -43,7 +43,7 @@ jobs:
4343
tests:
4444
name: Tests (Python ${{ matrix.python-version }})
4545
runs-on: ubuntu-latest
46-
timeout-minutes: 15
46+
timeout-minutes: 25
4747
strategy:
4848
fail-fast: false
4949
matrix:
@@ -54,9 +54,10 @@ jobs:
5454
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
5555

5656
- name: Install uv
57-
uses: astral-sh/setup-uv@37802adc94f370d6bfd71619e3f0bf239e1f3b78 # v7
57+
uses: astral-sh/setup-uv@c771a70e6277c0a99b617c7a806ffedaca235ff9 # v9.0.0
5858
with:
5959
enable-cache: true
60+
prune-cache: true
6061

6162
- name: Set up Python
6263
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7

.github/workflows/release-slow.yml

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -22,9 +22,10 @@ jobs:
2222
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
2323

2424
- name: Install uv
25-
uses: astral-sh/setup-uv@37802adc94f370d6bfd71619e3f0bf239e1f3b78 # v7
25+
uses: astral-sh/setup-uv@c771a70e6277c0a99b617c7a806ffedaca235ff9 # v9.0.0
2626
with:
2727
enable-cache: true
28+
prune-cache: true
2829

2930
- name: Set up Python
3031
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7

CHANGELOG.md

Lines changed: 94 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,100 @@ Detailed development history for the EvidenceForge project. Transferred from TOD
66

77
## Unreleased
88

9+
## v1.15.0 (2026-08-10)
10+
11+
This minor release completes the non-breaking canonical event and action-bundle
12+
architecture, substantially strengthens cross-source lifecycle realism, and
13+
incorporates repeated blind-assessment fixes across endpoint, network, email,
14+
and evaluation behavior.
15+
16+
**Canonical architecture and safety**
17+
18+
- Added the canonical contract foundation and completed occurrence, dispatch,
19+
session/authentication, interactive-session, network-projection,
20+
world-capability, evaluation-validity, and workload-safety boundaries
21+
(`0b0ad32c`, `b89855c9`, `739dbea2`, `b2f1dd68`, `91247365`, `c139c77a`,
22+
`03401ad0`, `868eb35d`, `53934a16`).
23+
- Closed the post-migration realism blockers without restoring compatibility
24+
layers or changing the public scenario schema (`2ab8e5e9`).
25+
26+
**Windows endpoint, identity, and session realism**
27+
28+
- Enforced module lifecycles, channel-time-derived record IDs, and more varied
29+
startup module profiles (`49148992`, `f4979524`, `df33ad10`).
30+
- Modeled authentication attempts and privilege-event ownership, and kept
31+
desktop bootstrap and session lifecycles attached to their canonical logons
32+
(`cfd6e10a`, `5c7786d6`, `f5604488`, `2942320d`).
33+
- Made enterprise software cohorts and endpoint singleton planning coherent
34+
and atomic across process and network ownership (`c0a01395`, `b96a96be`,
35+
`4a816bcc`, `b7275437`, `d5c6900a`).
36+
- Reused resident desktop resources and named service processes, and bound
37+
endpoint effects and ambient Windows file activity to native process owners
38+
(`b3cd49a0`, `d7162d57`, `4f062448`, `90a997db`).
39+
- Preserved host timezone state and canonical shell bootstrap timing, modeled
40+
complete Windows session-process lifecycles, and prevented duplicate RDP
41+
shell bootstrap (`bfbe3ed3`, `a1227098`, `ea914c19`, `81fc795c`, `4cc75358`).
42+
43+
**Remote access and Linux lifecycle ownership**
44+
45+
- Enforced Linux process-observation chronology and bounded SSH client
46+
ownership, then attached Windows SSH clients, credentials, receiver
47+
lifecycles, routine scheduling, and responder processes to their owning
48+
sessions (`8616105d`, `a8ddb632`, `2312ec24`, `eeb97e11`, `00a36029`,
49+
`316f0da7`, `84294984`).
50+
- Preserved explicit-credential source semantics and modeled one-shot caller
51+
behavior, while assigning Linux SMB activity to native client processes
52+
(`f0c6117d`, `17f73d05`, `ebab5c8f`).
53+
54+
**Network, protocol, proxy, and source-native projection**
55+
56+
- Ordered HTTP activity and made persistent transactions share coherent,
57+
durable observation identities (`04eb354c`, `8f1593f3`, `9dc294ad`,
58+
`481f0ed8`).
59+
- Ordered and retained ASA translations, modeled coherent sensor clocks and
60+
passive datagram accounting, and clipped perimeter teardown observations to
61+
visible intervals (`20f08d56`, `2ed74a5d`, `dc4d519c`, `0e75af19`,
62+
`e5f5aac8`).
63+
- Defined and annotated proxy tunnel accounting, omitted unavailable totals,
64+
reconciled byte scopes, and attached proxy flows to native client processes
65+
(`76421d75`, `1f9e4f5d`, `1f854be8`, `998d9942`, `1ee7ad95`, `ea53038a`).
66+
- Enforced half-open collection windows, gated analyzers on captured content,
67+
normalized Zeek histories, preserved TLS certificate observation, and
68+
excluded implausible external client networks (`4565b2e0`, `9e3e7c28`,
69+
`e3e9ed9a`, `fc51a85d`, `1f9dd5d1`).
70+
- Rendered source-native output against explicit versioned schemas
71+
(`c3f17ead`).
72+
73+
**Long-lived services, maintenance, packages, and applications**
74+
75+
- Honored bounded process deadlines and completed ownership/closure for APT
76+
frontends, package managers, and their transport helpers (`318dfe59`,
77+
`490e363b`, `eba43cd9`, `ed2e4d80`).
78+
- Modeled durable ServiceHealth agents and stateful fleet maintenance
79+
(`2a205e0b`, `294d2c80`).
80+
- Reused resident mail clients, preserved durable but varied rsyslog health
81+
state, and aligned generated email artifacts with their owning applications
82+
(`07751e47`, `481a48d1`, `772ebabc`, `60ae2587`).
83+
84+
**Evaluation correctness and regression coverage**
85+
86+
- Reconciled storyline pivot identity, explicit-proxy paths, duration windows,
87+
account deletion, and truthful already-satisfied action no-ops, restoring
88+
full pivot linkability on the expanded IDS assessment (`dcb89ca6`).
89+
- Corrected the private-IPv6 network regression expectation (`f17096f0`).
90+
91+
**Assessment, documentation, and CI**
92+
93+
- Published the complete realism review, reconciled its remediation roadmap,
94+
and recorded the post-Batch-2 and post-Batch-7b effectiveness gates
95+
(`7238ca61`, `2b473a3d`, `f2a15f6a`, `0934e808`, `ca03d731`).
96+
- Recorded assessment loops 15-34 and their individual handoffs, then completed
97+
the expanded iteration-test loop documentation (`ff7928e0`, `2eb98588`,
98+
`53c454aa`, `6472d02a`, `45467a6b`, `161f30ff`, `6339e586`).
99+
- Extended CI time allowance for the expanded validation suite (`40ac5f71`).
100+
- Updated Ruff to 0.16.1 and setup-uv to 9.0.0, retaining explicit cache
101+
pruning for cache-enabled GitHub Actions jobs (`8c4685c7`, `93a4510e`).
102+
9103
## v1.14.1 (2026-08-04)
10104

11105
This patch release corrects source-native network and endpoint behavior found

README.md

Lines changed: 29 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -13,7 +13,11 @@ For background on the project and why we built it, read our announcement:
1313

1414
Most synthetic log generators produce isolated, single-format data that experienced analysts identify as fake within seconds. EvidenceForge takes a fundamentally different approach:
1515

16-
- **Consistency by construction.** A canonical `SecurityEvent` model feeds all log formats from a single source of truth. Two emitters cannot disagree about a port number, timestamp, or LogonID because there is only one value — on the event object. This eliminates the cross-source inconsistencies that are the #1 tell of synthetic data.
16+
- **Consistency by construction.** Canonical events and immutable plans own shared connection,
17+
identity, lifecycle, and protocol facts for migrated action families; emitters perform only
18+
source-native projection. Dispatch rejects unknown or incomplete canonical occurrences. Some
19+
mutable compatibility views remain during the approved migration, and the explicit `raw` path
20+
is outside cross-source consistency guarantees.
1721

1822
- **Causal event ordering.** Events respect real-world dependencies — DNS queries precede connections, Kerberos TGT/TGS precede domain logons, audit events follow administrative commands. A composable rule engine auto-generates prerequisites with realistic timing offsets, so the data tells a coherent causal story across log sources.
1923

@@ -25,7 +29,9 @@ Most synthetic log generators produce isolated, single-format data that experien
2529

2630
- **Deterministic engine, LLM-assisted authoring.** Scenario creation uses Claude Code Skills for interactive, research-backed attack planning. Log generation is fully deterministic — no LLM calls, no API costs, reproducible output every time.
2731

28-
- **Built-in quality evaluation.** A 4-pillar scoring framework (21 sub-scores) measures parseability, plausibility, causality, and timing. Know exactly how good your data is before using it.
32+
- **Built-in quality evaluation.** A 4-pillar scoring framework (22 sub-scores) measures
33+
parseability, plausibility, causality, and timing, with additional concern-oriented views for
34+
source schema, canonical invariants, scenario completeness, and distribution realism.
2935

3036
## Quick Start
3137

@@ -80,9 +86,9 @@ For scripted or non-interactive use:
8086

8187
| Command | Description |
8288
|---------|-------------|
83-
| `eforge generate <scenario.yaml> -o <dir>` | Generate logs from a scenario file |
84-
| `eforge validate <scenario.yaml>` | Validate scenario schema and cross-references |
85-
| `eforge eval <output_dir> -s <scenario.yaml>` | Evaluate data quality (4 pillars, 21 sub-scores) |
89+
| `eforge generate <scenario.yaml> -o <dir> [--seed N] [--allow-large-workload]` | Generate logs; `--seed` overrides the scenario seed and the trusted workload override bypasses resource limits, not path safety |
90+
| `eforge validate <scenario.yaml> [--allow-large-workload]` | Validate schema, cross-references, and the default workload envelope |
91+
| `eforge eval <output_dir> -s <scenario.yaml> [--allow-large-evaluation]` | Evaluate quality (4 pillars, 22 sub-scores); the trusted override bypasses evaluator corpus limits |
8692
| `eforge info [field]` | Show installation info, config paths, and data inventories. Pass a dot-path field for a specific value (e.g., `eforge info personas`). Use `--fields` to list available fields, `--json` for machine output. |
8793
| `eforge validate-config` | Validate config files for cross-reference integrity. Use `--json` for machine output. |
8894
| `eforge install-skills [--agent all\|claude\|chatgpt\|codex] [--global]` | Install project-local or user-wide agent skills; defaults to all agents (`codex` aliases `chatgpt`) |
@@ -128,7 +134,8 @@ Every generated scenario includes a `GROUND_TRUTH.md` file. Attack scenarios doc
128134
- **Ground truth documentation** — Every run generates a GROUND_TRUTH.md; attack scenarios include narrative, timeline, and IOCs
129135
- **Parallel generation** — Threaded emitters write all formats simultaneously with temporal consistency
130136
- **Scenario validation** — Cross-reference checking, uniqueness constraints, and network topology validation
131-
- **Data quality evaluation** — 4-pillar scoring framework (21 sub-scores) with acceptance criteria
137+
- **Data quality evaluation** — 4-pillar scoring framework (22 sub-scores), concern-oriented
138+
diagnostic categories, non-vacuous applicability, and acceptance criteria
132139
- **Multi-timezone support** — Pattern-based timezone overrides per system hostname
133140

134141
## Supported Log Formats
@@ -214,7 +221,19 @@ EvidenceForge includes a built-in evaluation framework that scores generated dat
214221
| Causality | 25% | Causal ordering, event presence, indicator accuracy, pivot linkability |
215222
| Timing | 20% | Attack-chain timing, burstiness, diurnal patterns, volume adequacy |
216223
217-
**Two-tier acceptance**: hard gates (minimum, must pass) + aspirational targets (stretch goals, informational). Hard gates: Spec Conformance ≥ 95%, Value Plausibility ≥ 95%, Causal Ordering ≥ 90%, Event Presence ≥ 85%. Thresholds are configurable in `src/evidenceforge/config/evaluation/thresholds.yaml`.
224+
**Two-tier acceptance**: applicable hard gates must pass, while aspirational targets remain
225+
informational. Gates cover source conformance, value and field consistency, exact IDS integrity,
226+
causal/scenario reconciliation, and linkability. The authoritative thresholds are configurable in
227+
`src/evidenceforge/config/evaluation/thresholds.yaml`.
228+
229+
The compatibility pillars remain the weighted public score. The report also groups the same
230+
applicable measures by review concern: source/schema fidelity, canonical cross-source invariants,
231+
declared-scenario completeness, and distribution realism. Required measures with no applicable
232+
denominator are reported as unavailable instead of receiving a vacuous perfect score.
233+
234+
By default, evaluation accepts at most 512 MiB of input, 10,000 files, and 500,000 parsed records.
235+
Use `--allow-large-evaluation` only for a trusted corpus after reviewing available memory. The
236+
override does not relax file or parser safety checks.
218237

219238
```bash
220239
uv run eforge eval ./output -s scenario.yaml
@@ -235,7 +254,7 @@ GenerationEngine (hour-by-hour orchestration)
235254
WorldModel / WorldPlanner (compile host roles, user placement, session bootstrap)
236255
|
237256
v
238-
ActivityGenerator (builds SecurityEvents with composable contexts)
257+
ActivityGenerator (builds canonical occurrences with composable contexts)
239258
|
240259
v
241260
EventDispatcher (routes to StateManager + matching emitters)
@@ -257,7 +276,7 @@ emitters apply it only where file shape differs.
257276
258277
`WorldModel` compiles authoritative host and user capabilities from scenario fields like `primary_system`, `roles`, `services`, and workstation assignments. `WorldPlanner` then chooses realistic interactive, network, SSH, and RDP session paths before `ActivityGenerator` emits the correlated evidence.
259278
260-
See [Architecture Documentation](docs/ARCHITECTURE.md) for the full deep dive including the world-model layer, SecurityEvent model, state management, and emitter system.
279+
See [Architecture Documentation](docs/ARCHITECTURE.md) for the full deep dive including the world-model layer, CanonicalOccurrence model, state management, and emitter system.
261280
262281
## Development
263282
@@ -312,7 +331,7 @@ full-dataset runner command, and failure report details.
312331
### Design Documents
313332

314333
- [PRD](docs/design/PRD.md) — Product requirements and specifications
315-
- [Event Model Design](docs/design/event-model-prd.md) — Canonical SecurityEvent architecture
334+
- [Event Model Design](docs/design/event-model-prd.md) — Canonical occurrence architecture
316335
- [Data Quality Design](docs/design/data-quality-prd.md) — Evaluation framework design
317336
- [Research Report](docs/design/synthetic-log-generation-research.md) — Analysis of existing tools
318337

0 commit comments

Comments
 (0)